EDBT 2026 Demo / reviewers in the wild / expert
Yu Duan 0001
dblp:182/9655-1
· DBLP profile ↗
22ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0003-1799-7698ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Intrinsic Hierarchy for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) aims to classify unlabeled data by leveraging knowledge from labeled categories. While existing methods have achieved remarkable progress, they often treat images as flat feature sets, neglecting the intrinsic hierarchy: where key objects dominate meaning and backgrounds serve as context. For instance, in images of a dog either standing on grass or lying on a bed, the dog remains the central semantic element, whereas the background varies. Motivated by this, we propose LEArning Intrinsic Hierarchy (LEAH), a lightweight plug-and-play module designed to model hierarchical structure within images. LEAH consists of two components: a pruner that filters task-irrelevant tokens to extract key objects, and a constructor that embeds key objects and full images into hyperbolic space using adaptive entailment cones to capture compositional semantics. LEAH can be easily integrated into existing GCD frameworks with minimal modification. When applied to SimGCD, it achieves up to 13.2% accuracy improvement on fine-grained benchmarks, demonstrating its effectiveness in discovering subtle inter-class differences through hierarchical modeling. Yu Duan 0001, Junzhi He, Zhanxuan Hu, Mengda Ji, Rong Wang 0001, Quanxue Gao |
AAAI | 1 |
| 2026 | Maximizing Schatten-p Norm Regularization Toward BalanceabstractThe Schatten-p norm, as a class of structure-inducing norms based on singular values, has been widely used to enhance model low-rankness and representation capability due to its flexibility in structural modeling and favorable mathematical properties. However, its potential in cluster distribution modeling has long been overlooked. Therefore, we explore the potential of maximizing the Schatten-p norm as a regularization strategy specifically designed to achieve balanced clustering. This work is the first to investigate its effectiveness in promoting cluster balance. To be specific, maximizing Schatten-p norm effectively guides the assignment of data points, ensuring a more balanced distribution of samples across clusters. We have conducted an in-depth theoretical analysis and validated its effectiveness through extensive clustering experiments. Experimental results demonstrate that, compared to existing methods, this regularization term significantly improves clustering quality and obtain reasonable clustering. Fangfang Li 0005, Quanxue Gao, Yu Duan 0001, Yuzhuo Feng, Qin Li 0001 |
AAAI | 4 |
| 2026 | Balanced federated multi-view clustering with one-round communication
Quanxue Gao, Yu Duan 0001 |
Neurocomputing | 3 |
| 2026 | Probabilistic regression-based multi-view clustering with anchor graphs
Qin Li 0001, Quanxue Gao, Yu Duan 0001, Cheng Deng 0002 |
Neurocomputing | 5 |
| 2026 | Discriminative and noise-robust embedding for zero-shot learning
Cheng Deng 0002, Yu Duan 0001, Quanxue Gao |
Neural Networks | 3 |
| 2026 | Anchor-Guided Discrete Multi-View ClusteringabstractMulti-view clustering based on anchor graphs has attracted significant attention due to its ability to substantially reduce computational complexity, enabling the efficient processing of large-scale multimedia data. However, most existing anchor graph-based clustering methods fail to fully exploit the intrinsic properties of anchor graphs when applying regression techniques. Moreover, some approaches focus solely on sample labels while overlooking the crucial role of anchor labels in clustering. To address these limitations, we leverage the probabilistic information of the anchor graph by employing probabilistic projection to map the anchor graph into the label space, thereby obtaining anchor labels. By clustering both anchors and samples simultaneously, the anchor graph serves as a guide to induce anchor labels, which are then used to generate sample labels, facilitating anchor-guided sample clustering. Furthermore, we propose a novel regularization paradigm based on the matrix nuclear norm, ensuring that the obtained results remain discrete and that the sample distribution across clusters is balanced. Additionally, we introduce a new matrix nuclear norm optimization method based on the first-order Taylor expansion. Extensive experiments on real-world datasets demonstrate the effectiveness and robustness of our proposed method, achieving superior performance compared to state-of-the-art approaches. Our code is available at:https://github.com/harunakai/ADMC Ran Jing, Quanxue Gao, Yu Duan 0001, Cheng Deng 0002, Ming Yang 0024 |
IEEE Trans. Multim. | 3 |
| 2025 | A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionabstractGeneralized Category Discovery (GCD) aims to classify unlabeled data from both known and unknown categories by leveraging knowledge from labeled known categories. While existing methods have made notable progress, they often overlook a hidden stumbling block in GCD: distracted attention. Specifically, when processing unlabeled data, models tend to focus not only on key objects in the image but also on task-irrelevant background regions, leading to suboptimal feature extraction. To remove this stumbling block, we propose Attention Focusing (AF), an adaptive mechanism designed to sharpen the model's focus by pruning non-informative tokens. AF consists of two simple yet effective components: Token Importance Measurement (TIME) and Token Adaptive Pruning (TAP), working in a cascade. TIME quantifies token importance across multiple scales, while TAP prunes non-informative tokens by utilizing the multi-scale importance scores provided by TIME. AF is a lightweight, plug-and-play module that integrates seamlessly into existing GCD methods with minimal computational overhead. When incorporated into one prominent GCD method, SimGCD, AF achieves up to 15.4% performance improvement over the baseline with minimal computational overhead. The implementation code is provided in https://github.com/Afleve/AFGCD. Qiyu Xu, Zhanxuan Hu, Yu Duan 0001, Ercheng Pei, Yonghang Tai |
ICCV | 3 |
| 2025 | High-Quality Label Learning in Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is a recently proposed open-world problem that aims to automatically classify and discover new categories based on partially labeled data. For unlabeled data, previous research commonly considers using pseudo-labels to assist in model learning. These pseudo-labels, together with the true labels of labeled data, form the learning targets for the final classifier, leading to better predictive outcomes. However, low-quality labels can inevitably hinder the learning process of the model. To address this issue, inspired by previous methods, we propose the Calibrated Generalized Category Discovery (CGCD) framework, which incorporates a projection head, a classifier head, and a calibration head. The projection head is used for representation learning. The calibration head learns high-quality labels from the robust predictions of the classifier head, and the classifier head utilizes these high-quality labels for more efficient learning. Both heads mutually enhance each other during training, ultimately leading to a superior solution. In addition, leveraging the characteristics of both the classifier head and calibration head, we designed a classifier representation distribution regularization term to further ensure consistency in their learning processes. Extensive experimental results demonstrate that the proposed CGCD framework achieves state-of-the-art performance across five general and fine-grained visual recognition datasets by leveraging high-quality label learning. Yu Duan 0001, Junzhi He, Feiping Nie 0001, Quanxue Gao, Cheng Deng 0002 |
ICDM | 1 |
| 2025 | Balanced and Discrete Projection Collaborative Clustering with Probabilistic RegressionabstractAnchor graph-based multi-view clustering significantly reduces computational complexity, but most existing methods still exhibit the following drawbacks: 1. They neglect the probabilistic nature of the anchor graph. 2. They focus solely on the sample label matrix and fail to account for the relationship between the sample labels and the anchor labels. To address these issues, this paper proposes an balanced and discrete multi-view clustering model based on concept probabilistic regression. Specifically, according to the probabilistic characteristics of the anchor graph, we summarize the relationships between the anchor graph, anchor labels, and sample labels, and propose a probabilistic regression model that better guides the learning of sample labels by imposing constraints on the anchor labels. Moreover, unlike the widely used$\ell_{2,1}$-norm minimization methods in machine learning, we maximize the$\ell_{2,1}$-norm to ensure a balanced distribution of anchors within each clustering. Experimental results on several publicly available datasets demonstrate that the proposed model achieves superior clustering performance compared to mainstream algorithms. Yu Duan 0001, Quanxue Gao, Xuhong Dong |
ICDM | 2 |
| 2025 | Multi-view Clustering Based on Probabilistic Tensor RegressionabstractMulti-view clustering based on anchor graph and regression is widely used to deal with high dimensional and redundant data. However, most of these methods ignore the probabilistic characteristics of anchor graph, and the effective information in different views is not fully mined. To solve these problems, we propose a multi-view clustering method based on probabilistic tensor regression (MVCPTR). Specifically, we reinterpret the regression process of the anchor graph from the perspective of probability. By modeling the anchor graph as the transition probability from samples to anchors, we construct the implicit relationship between labels of samples and anchors. In order to further mine the complementary information of multi-view data, we extend the anchor graph matrix regression to tensor regression to achieve multi-level information fusion at the representational level and decision level, and impose the Schatten p-norm constraint on the anchor label tensor and the sample label tensor to realize the bi-clustering of the anchors and samples. A large number of experiments prove the effectiveness of our proposed algorithm. Yichen Bao, Yu Duan 0001, Jing Li 0026, Quanxue Gao |
ACM Multimedia | 3 |
| 2025 | Fast multi-view clustering via anchor label transmit with tensor structure constraint
Runxin Zhang, Yu Duan 0001, Rong Wang 0001, Feiping Nie 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Prototypical classifier with distribution consistency regularization for generalized category discovery: A strong baseline
Zhanxuan Hu, Yu Duan 0001, Yaming Zhang, Rong Wang 0001, Feiping Nie 0001 |
Neural Networks | 2 |
| 2025 | Soft Neighbors Supported Contrastive ClusteringabstractExisting deep clustering methods leverage contrastive or non-contrastive learning to facilitate downstream tasks. Most contrastive-based methods typically learn representations by comparing positive pairs (two views of the same sample) against negative pairs (views of different samples). However, we spot that this hard treatment of samples ignores inter-sample relationships, leading to class collisions and degrade clustering performances. In this paper, we propose a soft neighbor supported contrastive clustering method to address this issue. Specifically, we first introduce a concept called perception radius to quantify similarity confidence between a sample and its neighbors. Based on this insight, we design a two-level soft neighbor loss that captures both local and global neighborhood relationships. Additionally, a cluster-level loss enforces compact and well-separated cluster distributions. Finally, we conduct a pseudo-label refinement strategy to mitigate false negative samples. Extensive experiments on benchmark datasets demonstrate the superiority of our method. The code is available at https://github.com/DuannYu/soft-neighbors-supported-clustering. Yu Duan 0001, Runxin Zhang, Rong Wang 0001, Feiping Nie 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Unconstrained Fuzzy C-Means Based on Entropy Regularization: An Equivalent ModelabstractFuzzy c-means based on entropy regularization (FCER) is a commonly used machine learning algorithm that uses maximum entropy as the regularization term to realize fuzzy clustering. However, this model has many constraints and is challenging to optimize directly. During the solution process, the membership matrix and cluster centers are alternately optimized, easily converging to poor local solutions, limiting the clustering performance. In this paper, we start with the optimization model and propose an unconstrained fuzzy clustering model (UFCER) equivalent to FCER, which reduces the size of optimization variables from$(n+d)\times c$to$d\times c$. More importantly, there is no need to calculate the membership matrix during the optimization process iteratively. The time complexity is only linear, and the convergence speed is fast. We conduct extensive experiments on real datasets. The comparison of objective function value and clustering performance fully demonstrates that under the same initialization, our proposed algorithm can converge to smaller local minimums and get better clustering performance. Feiping Nie 0001, Runxin Zhang, Yu Duan 0001, Rong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Toward Balance Deep Semisupervised ClusteringabstractThe goal of balanced clustering is partitioning data into distinct groups of equal size. Previous studies have attempted to address this problem by designing balanced regularizers or utilizing conventional clustering methods. However, these methods often rely solely on classic methods, which limits their performance and primarily focuses on low-dimensional data. Although neural networks exhibit effective performance on high-dimensional datasets, they struggle to effectively leverage prior knowledge for clustering with a balanced tendency. To overcome the above limitations, we propose deep semisupervised balanced clustering, which simultaneously learns clustering and generates balance-favorable representations. Our model is based on the autoencoder paradigm incorporating a semisupervised module. Specifically, we introduce a balance-oriented clustering loss and incorporate pairwise constraints into the penalty term as a pluggable module using the Lagrangian multiplier method. Theoretically, we ensure that the proposed model maintains a balanced orientation and provides a comprehensive optimization process. Empirically, we conducted extensive experiments on four datasets to demonstrate significant improvements in clustering performance and balanced measurements. Our code is available at https://github.com/DuannYu/BalancedSemi-TNNLS. Yu Duan 0001, Zhoumin Lu, Rong Wang 0001, Xuelong Li 0001, Feiping Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Hyperbolic Hierarchical Representation Learning for Generalized Category DiscoveryabstractThis study addresses the problem of generalized category discovery (GCD), an advanced and challenging semi-supervised learning scenario that deals with unlabeled data from both known and novel categories. Although recent research has effectively engaged with this issue, these studies typically map features into Euclidean space, which fails to maintain the latent semantic hierarchy of the training samples effectively. This limitation restricts the exploration of more detailed and rich information and degrades the performance in discovering new categories. The emerging field of hyperbolic representation learning suggests that hyperbolic geometry could be advantageous for extracting semantic information to tackle this problem. Motivated by this, we proposed hyperbolic hierarchical representation learning for GCD (HypGCD). Specifically, HypGCD enhances representations in hyperbolic space, building upon the Euclidean space representation from two perspectives: instance-class level and instance-instance level. At the instance-class level, HypGCD endeavors to construct well-defined clusters, with each sample forming a robust hierarchical cluster structure. Concurrently, at the instance-instance level, HypGCD anticipates that a subset of samples will display a tree-like structure in local space, which aligns more closely with real-world scenarios. Finally, HypGCD optimizes the Euclidean and hyperbolic space collectively to obtain refined features. Additionally, we show that HypGCD is exceptionally effective, achieving state-of-the-art (SOTA) results on several datasets. The code is available at https://github.com/DuannYu/HypGCD. Yu Duan 0001, Feiping Nie 0001, Zhanxuan Hu, Rong Wang 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | New approach for learning structured graph with Laplacian rank constraint
Yu Duan 0001, Feiping Nie 0001, Rong Wang 0001, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2024 | Harmonic cut: An efficient and directly solved balanced graph clustering
Yu Duan 0001, Feiping Nie 0001, Rong Wang 0001, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2024 | Scalable and parameter-free fusion graph learning for multi-view clustering
Yu Duan 0001, Danyang Wu, Rong Wang 0001, Xuelong Li 0001, Feiping Nie 0001 |
Neurocomputing | 1 |
| 2024 | MaskRecon: High-quality human reconstruction via masked autoencoders using a single RGB-D image
Xing Li 0040, Yangyu Fan, Zhibo Rao, Yu Duan 0001, Shiya Liu |
Neurocomputing | 5 |
| 2024 | Fuzzy Clustering From Subset-Clustering to Fullset-MembershipabstractFuzzy theory, which extends precise binary logic to continuous fuzzy logic, provides an effective tool for uncertainty problems in machine learning and thus, evolved fuzzy methods such as fuzzy c-means and fuzzy graph clustering. Among them, graph-based clustering methods have become a hot spot in the field of unsupervised clustering due to their good ability to process nonlinear data. Unfortunately, they usually suffer from high time complexity and cumbersome regularization parameter tuning, so their practical applications are greatly limited. To this end, we propose an efficient graph-cut algorithm called fS2F. Based on the similarity graph between the dataset and landmark subset, fS2F transforms the membership learning of the entire dataset into a clustering problem of representative points, which greatly improves its clustering efficiency. In addition, fS2F softly constrains the cluster size in a way that does not require additional regularization parameters so that it can be widely and conveniently applied. The article also presents the optimization method for this model and demonstrates its effectiveness through experiments. Yu Duan 0001, Feiping Nie 0001, Rong Wang 0001, Xuelong Li 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Unsupervised Deep Embedding for Fuzzy ClusteringabstractDeep fuzzy clustering employs neural networks to discover the low-dimensional embedding space of data, providing an effective solution to the clustering problem posed by high-dimensional data. Although some algorithms have achieved good results in application, the field still faces the following problems: the lack of clustering loss function limits the development of deep clustering, and most of them use the self-training strategy-based Kullback-Leibler (KL) divergence; some algorithms directly use conventional constrained clustering objective function as the loss function in deep models, and update network parameters alternately, the optimization process is cumbersome. Focusing on the issues mentioned above, this article first proposed an unconstrained fuzzy$c$-means algorithm that can be solved using gradient descent and then used it as the clustering loss function to obtain a novel deep fuzzy clustering model named unsupervised deep embedding for fuzzy clustering. The proposed model simultaneously learns the low-dimensional representation of data and performs fuzzy clustering. It updates parameters through gradient descent and backpropagation, achieving end-to-end optimization. The proposed algorithm's effectiveness and competitiveness are fully demonstrated through extensive experiments conducted on image and text datasets. Runxin Zhang, Yu Duan 0001, Feiping Nie 0001, Rong Wang 0001, Xuelong Li 0001 |
IEEE Trans. Fuzzy Syst. | 2 |