VLDB 2026 Research / reviewers in the wild / expert
Yu Wang 0106
dblp:02/5889-106
· DBLP profile ↗
53ranked-venue papers
9as first author
48since 2021 · last 2026
0000-0002-4788-8655ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 8 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Incomplete cross-modality class-incremental learning in visible-thermal recognitionabstractVisible-thermal cross-modality learning enhances downstream task performance by integrating information from multiple sources. In real-world scenarios such as autonomous driving, new classes continually emerge, and data is often incomplete due to sensor occlusions. This raises a key question about how to incrementally update a model with incomplete cross-modality data. To address this problem, we propose a practical task termed incomplete cross-modality class-incremental learning (ICMCIL), which aims to effectively leverage incomplete cross-modality information to learn new knowledge without forgetting the old. We construct a benchmark for ICMCIL and thoroughly analyze its challenges, revealing that (1) different modalities experience varying degrees of forgetting, (2) conventional cross-modality fusion only partially alleviates forgetting, and (3) missing data exacerbates the forgetting of previous classes. To address these issues, we propose Hybrid Fusion via Completion (HFC), a unified framework that integrates completion, fusion, and forgetting prevention. Additionally, we enhance information fusion by introducing a feature interchange mechanism, wherein features are shuffled and channels are reordered to improve information flow. Extensive experiments demonstrate that HFC effectively addresses ICMCIL, significantly mitigating modality forgetting. Xinjie Yao, Yanxian Bi, Yu Wang 0106, Pengfei Zhu 0001, Ruipu Zhao, Wanyu Lin, Qinghua Hu |
Pattern Recognit. | 3 |
| 2026 | FCGNN: Fuzzy Cognitive Graph Neural Networks With Concept Evolution for Few-Shot LearningabstractGraph neural networks (GNNs) have recently emerged as a promising approach for solving few-shot learning (FSL) problems, enabling generalization to new categories by establishing associations among limited samples. However, most previous methods focus on local interactions within a layer or between successive layers, overlooking the continuous evolution of node states across deeper layers. Especially in FSL scenarios with scarce data, this oversight leads to excessive compression or loss of node features in deep networks. To address this challenge, we propose a Fuzzy Cognitive Graph Neural Network (FCGNN), which models the evolution of node fuzzy state representation across layers by integrating principles from fuzzy cognitive maps. First, we construct a fuzzy cognitive graph (FCG) based on Gaussian similarity to enhance the information flow within the layer. FCG establishes fuzzy relationships between nodes to enhance the connections between different samples within the layer. Then, we design a fuzzy concept evolution (FCE) module based on dual fuzzy state representation, incorporating the fuzzy activation state to illustrate the continuous changes in node states. FCE combines explicit and implicit fuzzy state representation to achieve the dynamic evolution of node concepts between layers. Finally, we map the node states after multiple graph iterations to the category space for FSL image classification. In experiments on four benchmark datasets, FCGNN achieves superior classification performance and outperforms several baseline methods by up to 2.73% on average accuracy. Linhua Zou, Chengxi Jiang, Yu Wang 0106, Hong Zhao 0002 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2026 | Multi-Granularity Superpoint Graph Learning for Weakly Supervised 3D Semantic SegmentationabstractWeakly supervised 3D semantic segmentation has proven effective in alleviating the heavy dependence on dense annotations by generating high-quality pseudo-labels. However, due to the scene complexity and disorder of the point cloud, merely applying the model semantic prediction or hand-crafted feature similarity for pseudo labeling is inefficient and biased. This limitation inevitably results in incorrect pseudo labels. To tackle this challenge, we propose a new method called Multi-granularity Superpoint Graph Learning (MSGL) that leverages the multi-scale local features of point clouds to improve the quality of pseudo labels. We first design a multi-granularity local representation learning module on the superpoint graph to capture the neighboring structure information of each superpoint within complex scenes. Subsequently, the generated structural embedding is utilized to enhance the affinity matrix of label propagation, thereby yielding high-quality pseudo labels. To further enforce the generalization of the structural representation module under scenario changes or data fluctuations, we present a multi-granularity consistency loss in MSGL. This loss is applied across different views of the superpoint graph within each scene to ensure a robust and consistent learning process. Our experiments conducted on three benchmarks show that the proposed method outperforms existing weakly supervised methods under several sparse label settings, and improves the baseline by an average of 7.7% with only 1% extra computation cost. Moreover, our approach even compares favorably to some fully supervised methods with only one point labeled for each thing. Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Bin Xiao 0002, Qinghua Hu |
IEEE Trans. Multim. | 2 |
| 2025 | Reducing Class-wise Confusion for Incremental Learning with Disentangled ManifoldsabstractClass incremental learning (CIL) aims to enable models to continuously learn new classes without catastrophically forgetting old ones. A promising direction is to learn and use prototypes of classes during incremental updates. Despite simplicity and intuition, we find that such methods suffer from inadequate representation capability and unsatisfied feature overlap. These two factors cause class-wise confusion and limited performance. In this paper, we develop a Confusion-REduced AuTo-Encoder classifier (CREATE) for CIL. Specifically, our method employs a lightweight auto-encoder module to learn compact manifold for each class in the latent subspace, constraining samples to be well reconstructed only on the semantically correct auto-encoder. Thus, the representation stability and capability of class distributions are enhanced, alleviating the potential class-wise confusion problem. To further distinguish the overlapped features, we propose a confusion-aware latent space separation loss that ensures samples are closely distributed in their corresponding low-dimensional manifold while keeping away from the distributions of features from other classes. Our method demonstrates stronger representational capacity and discrimination ability by learning disentangled manifolds and reduces class confusion. Extensive experiments on multiple datasets and settings show that CREATE outperforms other state-of-the-art methods up to 5.41%. The code is available at https://github.com/lilyht/CREATE. Huitong Chen, Yu Wang 0106, Yan Fan 0002, Guosong Jiang, Qinghua Hu |
CVPR | 2 |
| 2025 | Long-Tailed Classification with Multi-Granularity Semantics
Liu Yang 0010, Yu Wang 0106 |
ICCV | 3 |
| 2025 | Socialized Coevolution: Advancing a Better World through Cross-Task CollaborationabstractTraditional machine societies rely on data-driven learning, overlooking interactions and limiting knowledge acquisition from model interplay. To address these issues, we revisit the development of machine societies by drawing inspiration from the evolutionary processes of human societies. Motivated by Social Learning (SL), this paper introduces a practical paradigm of Socialized Coevolution (SC). Compared to most existing methods focused on knowledge distillation and multi-task learning, our work addresses a more challenging problem: not only enhancing the capacity to solve new downstream tasks but also improving the performance of existing tasks through inter-model interactions. Inspired by cognitive science, we propose Dynamic Information Socialized Collaboration (DISC), which achieves SC through interactions between models specialized in different downstream tasks. Specifically, we introduce the dynamic hierarchical collaboration and dynamic selective collaboration modules to enable dynamic and effective interactions among models, allowing them to acquire knowledge from these interactions. Finally, we explore potential future applications of combining SL and SC, discuss open questions, and propose directions for future research, aiming to spark interest in this emerging and exciting interdisciplinary field. Our code will be publicly available at https://github.com/yxjdarren/SC. Xinjie Yao, Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Ruipu Zhao, Zhoupeng Guo, Qinghua Hu |
ICML | 2 |
| 2025 | Graphs Help Graphs: Multi-Agent Graph Socialized LearningabstractGraphs in the real world are fragmented and dynamic, lacking collaboration akin to that observed in human societies. Existing paradigms present collaborative information collapse and forgetting, making collaborative relationships poorly autonomous and interactive information insufficient. Moreover, collaborative information is prone to loss when the graph grows. Effective collaboration in heterogeneous dynamic graph environments becomes challenging. Inspired by social learning, this paper presents a Graph Socialized Learning (GSL) paradigm. We provide insights into graph socialization in GSL and boost the performance of agents through effective collaboration. It is crucial to determine with whom, what, and when to share and accumulate information for effective GSL. Thus, we propose the ''Graphs Help Graphs'' (GHG) method to solve these issues. Specifically, it uses a graph-driven organizational structure to select interacting agents and manage interaction strength autonomously. We produce customized synthetic graphs as an interactive medium based on the demand of agents, then apply the synthetic graphs to build prototypes in the life cycle to help select optimal parameters. We demonstrate the effectiveness of GHG in heterogeneous dynamic graphs by an extensive empirical study. The code is available through https://github.com/Jillian555/GHG. Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Xinjie Yao, Qinghua Hu |
NeurIPS | 2 |
| 2025 | Hyperbolic-Euclidean Deep Mutual LearningabstractGraph neural networks (GNNs) exhibit powerful performance in handling graph data, with Euclidean and hyperbolic variants excelling in processing grid-based and hierarchical structures, respectively. However, existing methods focus on learning specific structures linked to the inherent properties of the underlying space, failing to fully exploit their complementary properties in distinct geometric spaces, thus limiting their ability to efficiently model complex graph structures. In this paper, we propose a Hyperbolic-Euclidean Deep Mutual Learning (H-EDML) framework, which leverages the unique properties of hyperbolic space to effectively capture the hierarchical relationships present in graph data, while also utilizes the familiar Euclidean space to handle local interactions. Specifically, We design a topology mutual learning module to bolster the capacity of each single model to perceive the holistic topological structure of the graph. Then, we integrate a decision mutual learning module to further advance the models' comprehensive judgment capabilities towards graph data, thereby strengthening the robustness and generalization. Furthermore, we employ an attention-based probabilistic integration strategy for the final prediction to alleviate potential disparities in decision-making among different models. Extensive experiments on node classification are conducted on five real-world graph datasets and the results show that our proposed H-EDML achieves competitive performances compared to the state-of-the-art methods. The source code will be available at: https://github.com/caohaifang123/H-EDML. Haifang Cao, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu |
WWW | 2 |
| 2025 | Visible-thermal cross-modality class-incremental learning
Xinjie Yao, Yu Wang 0106, Pengfei Zhu 0001, Ruipu Zhao, Shenglei Pei, Wanyu Lin |
Expert Syst. Appl. | 3 |
| 2025 | BackMix: Regularizing Open Set Recognition by Removing Underlying Fore-Background PriorsabstractOpen set recognition (OSR) requires models to classify known samples while detecting unknown samples for real-world applications. Existing studies show impressive progress using unknown samples from auxiliary datasets to regularize OSR models, but they have proved to be sensitive to selecting such known outliers. In this paper, we discuss the aforementioned problem from a new perspective: Can we regularize OSR models without elaborately selecting auxiliary known outliers? We first empirically and theoretically explore the role of foregrounds and backgrounds in open set recognition and disclose that: 1) backgrounds that correlate with foregrounds would mislead the model and cause failures when encounters 'partially' known images; 2) Backgrounds unrelated to foregrounds can serve as auxiliary known outliers and provide regularization via global average pooling. Based on the above insights, we propose a new method, Background Mix (BackMix), that mixes the foreground of an image with different backgrounds to remove the underlying fore-background priors. Specifically, BackMix first estimates the foreground with class activation maps (CAMs), then randomly replaces image patches with backgrounds from other images to obtain mixed images for training. With backgrounds de-correlated from foregrounds, the open set recognition performance is significantly improved. The proposed method is quite simple to implement, requires no extra operation for inferences, and can be seamlessly integrated into almost all of the existing frameworks. Yu Wang 0106, Junxian Mu, Hongzhi Huang, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | FN-NET: Adaptive data augmentation network for fine-grained visual categorization
Shuo Ye, Qinmu Peng, Yiu-Ming Cheung, Yu Wang 0106, Ziqian Zou, Xinge You |
Pattern Recognit. | 4 |
| 2025 | Uncertainty-Aware Superpoint Graph Transformer for Weakly Supervised 3-D Semantic SegmentationabstractWeakly supervised 3-D semantic segmentation has successfully mitigated the labor-intensive and time-consuming task of annotating 3-D point clouds. However, reliably utilizing the minimal point-wise annotations for unlabeled data in complex and large-scale scenes is still challenging, such as only 20 points labeled in 2 million points. To tackle this challenge, we propose a new Uncertainty-aware Superpoint Graph Transformer (UaSGT) framework that utilizes minimal annotations for unlabeled data learning through reliable long-range supervision propagation from labeled superpoints to unlabeled superpoints. First, we propose a superpoint graph transformer to achieve long-range supervision propagation along the attention-based fuzzy subsets defined on superpoints. The attention-based fuzzy subset measures the membership of unlabeled superpoints to clusters centered on labeled superpoints. Second, we employ an uncertainty-aware membership rectification technique on the fuzzy subset to ensure reliable propagation among superpoints within the same category. This technique integrates an uncertainty prediction module to mask the influence of unreliable membership and a spatial prior refinement module to reduce uncertainty in intraclass membership degrees. Finally, experimental results on two large-scale benchmarks S3DIS and ScanNet-V2 demonstrate the superiority of our approach compared to the state-of-the-art with at least 90% annotation reduction, and our method also achieves comparable performance to fully supervised methods with less than 0.1% labeled points. Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Qinghua Hu |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | SPARK: Simple and Parameter-Free Knowledge Embedding With Fuzzy Cognitive Maps for Class Incremental LearningabstractClass incremental learning (CIL) aims to mitigate catastrophic forgetting of previously learned classes when integrating new knowledge. A primary challenge contributing to forgetting is the absence of data from earlier classes. Researchers have designed a variety of methods to solve the problem, among which topology-preserving methods show tremendous potential. However, two problems remain: (1) A large hyperparameter search space for constructing and utilizing a complex topology hinders efficient performance optimization; (2) Constraining the network to preserve the topology in the objective makes it difficult to optimize. This paper proposes SPARK, a simple and parameter-free method by embedding fuzzy cognitive maps, to address the problems. First, we construct a fuzzy cognitive map with nodes representing class prototypes and edges representing inter-class similarities. Then, we exploit the fuzzy cognitive map to obtain class-level embedding by aggregating features of other classes for each class. Finally, the class- and sample-level embeddings are fused and fed to the classifier. The proposed method can be easily optimized without introducing additional loss terms and hyperparameters. We theoretically prove that such a simple fuzzy cognitive map embedding can efficiently preserve the structural information of the fuzzy cognitive map. Experimental results indicate that SPARK achieves up to 5.85% higher average accuracy and 8.69% reduction in forgetting compared to baseline methods. The code is available athttps://github.com/helloxjb/SPARK/ Yu Wang 0106, Jiabo Xie, Junyan Zheng, Bingxu Lu, Yanxian Bi, Qinghua Hu |
IEEE Trans. Fuzzy Syst. | 1 |
| 2025 | CKD: Contrastive Knowledge Distillation From a Sample-Wise PerspectiveabstractIn this paper, we propose a simple yet effective contrastive knowledge distillation framework that achieves sample-wise logit alignment while preserving semantic consistency. Conventional knowledge distillation approaches exhibit over-reliance on feature similarity per sample, which risks overfitting, and contrastive approaches focus on inter-class discrimination at the expense of intra-sample semantic relationships. Our approach transfers "dark knowledge" through teacher-student contrastive alignment at the sample level. Specifically, our method first enforces intra-sample alignment by directly minimizing teacher-student logit discrepancies within individual samples. Then, we utilize inter-sample contrasts to preserve semantic dissimilarities across samples. By redefining positive pairs as aligned teacher-student logits from identical samples and negative pairs as cross-sample logit combinations, we reformulate these dual constraints into an InfoNCE loss framework, reducing computational complexity lower than sample squares while eliminating dependencies on temperature parameters and large batch sizes. We conduct comprehensive experiments across three benchmark datasets, including the CIFAR-100, ImageNet-1K, and MS COCO datasets, and experimental results clearly confirm the effectiveness of the proposed method on image classification, object detection, and instance segmentation tasks. Wencheng Zhu, Pengfei Zhu 0001, Yu Wang 0106, Qinghua Hu |
IEEE Trans. Image Process. | 4 |
| 2025 | Boosting Pseudo-Labeling With Curriculum Self-Reflection for Attributed Graph ClusteringabstractAttributed graph clustering is an unsupervised learning task that aims to partition various nodes of a graph into distinct groups. Existing approaches focus on devising diverse pretext tasks to obtain suitable supervised information for representation learning, among which the predictive methods show great potential. However, these methods 1) generate auxiliary task bias toward the clustering target and 2) introduce label noise due to static thresholds. To address this issue, we propose a new self-supervised learning method, namely, pseudo-labeling with curriculum self-reflection (PLCSR), that learns reliable pseudo-labels by mining its information to achieve progressive processing of nodes in a self-reflection manner. First, a self-auxiliary encoder is constructed using the exponential moving average (EMA) of the original encoder's parameters to replace the auxiliary tasks, which provides an additional perspective of finding highly confident pseudo-labels. Second, a curriculum selection strategy using dynamic thresholds is designed to take full advantage of graph nodes more accurately. Besides simple nodes with high confidence at the initial stage, nodes that yield consistent predictions from both encoders are then assigned pseudo-labels to avoid the under-learning problem. For the rest difficult nodes that are highly uncertain, we abstain from making judgments to minimize their adverse impact on the model. Extensive experiments have shown that PLCSR significantly outperforms the state-of-the-art predictive method CDRS, achieving more than 6% improvements in terms of clustering accuracy. The code is available at: https://github.com/Jillian555/PLCSR. Pengfei Zhu 0001, Yu Wang 0106, Bin Xiao 0002, Jinglin Zhang 0001, Wanyu Lin, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Multi-Granularity Weighted Federated Learning for Heterogeneous Edge ComputingabstractFederated learning (FL), an advanced variant of distributed machine learning, enables clients to collaboratively train a model without sharing raw data, thereby enhancing privacy, security, and reducing communication overhead. However, in edge computing scenarios, there is an increasing trend towards diversity, heterogeneity, and complexity in clients’ data and models. The fundamental challenges, such as non-independent and identically distributed (non-IID) data and multi-granularity data accompanied by model heterogeneity, have become more evident and pose challenges to collaborative training among clients. In this paper, we refine the FL framework and propose the Multi-granularity Weighted Federated Learning (MGW-FL), emphasizing efficient collaborative training among clients with varied data granularities and diverse model scales across distinct data distributions. We introduce a distance-based FL mechanism designed for homogeneous clients, providing personalized models to mitigate the negative effects that non-IID data might have on model aggregation. Simultaneously, we propose an attention-weighted FL mechanism enhanced by a prior attention mechanism, facilitating knowledge transfer across clients with heterogeneous data granularities and model scales. Furthermore, we provide theoretical analyses of the convergence properties of the proposed MGW-FL method for both convex and non-convex models. Experimental results on five benchmark datasets demonstrate that, compared to baseline methods, MGW-FL significantly improves accuracy by almost 150% and convergence efficiency by nearly 20% on both IID and non-IID data. Chao Qiu, Shangxuan Cai, Yu Wang 0106, Xiaofei Wang 0001, Qinghua Hu |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Fuzzy Trajectory Tracking Control of Under-Actuated Unmanned Surface Vehicles With Ocean Current and Input QuantizationabstractThis article focuses on the trajectory tracking control of under-actuated unmanned surface vehicles subject to unknown ocean current and input quantization. Regarding kinematics, we devise an extended-state-observer-based guidance law capable of compensating for ocean currents to track the intended trajectory. Concerning kinetics, we propose an event-triggered adaptive fuzzy quantization control law using a linear analytical model to depict input quantization, eliminating the need for prior quantization parameter information. A notable aspect is the reduction in both execution frequency and magnitude, thereby mitigating communication burdens. The stability of this control strategy is proofed through input-to-state stability analysis. Simulation experiments are conducted to affirm the viability of the event-triggered adaptive fuzzy quantization control strategy. Jun Ning, Yu Wang 0106, Eryue Wang, Lu Liu 0003, C. L. Philip Chen, Shaocheng Tong |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Every Node Is Different: Dynamically Fusing Self-Supervised Tasks for Attributed Graph ClusteringabstractAttributed graph clustering is an unsupervised task that partitions nodes into different groups. Self-supervised learning (SSL) shows great potential in handling this task, and some recent studies simultaneously learn multiple SSL tasks to further boost performance. Currently, different SSL tasks are assigned the same set of weights for all graph nodes. However, we observe that some graph nodes whose neighbors are in different groups require significantly different emphases on SSL tasks. In this paper, we propose to dynamically learn the weights of SSL tasks for different nodes and fuse the embeddings learned from different SSL tasks to boost performance. We design an innovative graph clustering approach, namely Dynamically Fusing Self-Supervised Learning (DyFSS). Specifically, DyFSS fuses features extracted from diverse SSL tasks using distinct weights derived from a gating network. To effectively learn the gating network, we design a dual-level self-supervised strategy that incorporates pseudo labels and the graph structure. Extensive experiments on five datasets show that DyFSS outperforms the state-of-the-art multi-task SSL methods by up to 8.66% on the accuracy metric. The code of DyFSS is available at: https://github.com/q086/DyFSS. Pengfei Zhu 0001, Yu Wang 0106, Qinghua Hu |
AAAI | 3 |
| 2024 | Dynamic Sub-graph Distillation for Robust Semi-supervised Continual LearningabstractContinual learning (CL) has shown promising results and comparable performance to learning at once in a fully supervised manner. However, CL strategies typically require a large number of labeled samples, making their real-life deployment challenging. In this work, we focus on semi-supervised continual learning (SSCL), where the model progressively learns from partially labeled data with unknown categories. We provide a comprehensive analysis of SSCL and demonstrate that unreliable distributions of unlabeled data lead to unstable training and refinement of the progressing stages. This problem severely impacts the performance of SSCL. To address the limitations, we propose a novel approach called Dynamic Sub-Graph Distillation (DSGD) for semi-supervised continual learning, which leverages both semantic and structural information to achieve more stable knowledge distillation on unlabeled data and exhibit robustness against distribution bias. Firstly, we formalize a general model of structural distillation and design a dynamic graph construction for the continual learning progress. Next, we define a structure distillation vector and design a dynamic sub-graph distillation algorithm, which enables end-to-end training and adaptability to scale up tasks. The entire proposed method is adaptable to various CL methods and supervision settings. Finally, experiments conducted on three datasets CIFAR10, CIFAR100, and ImageNet-100, with varying supervision ratios, demonstrate the effectiveness of our proposed approach in mitigating the catastrophic forgetting problem in semi-supervised continual learning scenarios. Our code is available: https://github.com/fanyan0411/DSGD. Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu |
AAAI | 2 |
| 2024 | Exploring Diverse Representations for Open Set RecognitionabstractOpen set recognition (OSR) requires the model to classify samples that belong to closed sets while rejecting unknown samples during test. Currently, generative models often perform better than discriminative models in OSR, but recent studies show that generative models may be computationally infeasible or unstable on complex tasks. In this paper, we provide insights into OSR and find that learning supplementary representations can theoretically reduce the open space risk. Based on the analysis, we propose a new model, namely Multi-Expert Diverse Attention Fusion (MEDAF), that learns diverse representations in a discriminative way. MEDAF consists of multiple experts that are learned with an attention diversity regularization term to ensure the attention maps are mutually different. The logits learned by each expert are adaptively fused and used to identify the unknowns through the score function. We show that the differences in attention maps can lead to diverse representations so that the fused representations can well handle the open space. Extensive experiments are conducted on standard and OSR large-scale benchmarks. Results show that the proposed discriminative method can outperform existing generative models by up to 9.5% on AUROC and achieve new state-of-the-art performance with little computational cost. Our method can also seamlessly integrate existing classification models. Code is available at https://github.com/Vanixxz/MEDAF. Yu Wang 0106, Junxian Mu, Pengfei Zhu 0001, Qinghua Hu |
AAAI | 1 |
| 2024 | Socialized Learning: Making Each Other Better Through Multi-Agent CollaborationabstractLearning new knowledge frequently occurs in our dynamically changing world, e.g., humans culturally evolve by continuously acquiring new abilities to sustain their survival, leveraging collective intelligence rather than a large number of individual attempts. The effective learning paradigm during cultural evolution is termed socialized learning (SL). Consequently, a straightforward question arises: Can multi-agent systems acquire more new abilities like humans? In contrast to most existing methods that address continual learning and multi-agent collaboration, our emphasis lies in a more challenging problem: we prioritize the knowledge in the original expert classes, and as we adeptly learn new ones, the accuracy in the original expert classes stays superior among all in a directional manner. Inspired by population genetics and cognitive science, leading to unique and complete development, we propose Multi-Agent Socialized Collaboration (MASC), which achieves SL through interactions among multiple agents. Specifically, we introduce collective collaboration and reciprocal altruism modules, organizing collaborative behaviors, promoting information sharing, and facilitating learning and knowledge interaction among individuals. We demonstrate the effectiveness of multi-agent collaboration in an extensive empirical study. Our code will be publicly available at https://github.com/yxjdarren/SL. Xinjie Yao, Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Qinghua Hu |
ICML | 2 |
| 2024 | Persistence Homology Distillation for Semi-supervised Continual LearningabstractSemi-supervised continual learning (SSCL) has attracted significant attention for addressing catastrophic forgetting in semi-supervised data. Knowledge distillation, which leverages data representation and pair-wise similarity, has shown significant potential in preserving information in SSCL. However, traditional distillation strategies often fail in unlabeled data with inaccurate or noisy information, limiting their efficiency in feature spaces undergoing substantial changes during continual learning. To address these limitations, we propose Persistence Homology Distillation (PsHD) to preserve intrinsic structural information that is insensitive to noise in semi-supervised continual learning. First, we capture the structural features using persistence homology by homological evolution across different scales in vision data, where the multi-scale characteristic established its stability under noise interference. Next, we propose a persistence homology distillation loss in SSCL and design an acceleration algorithm to reduce the computational cost of persistence homology in our module. Furthermore, we demonstrate the superior stability of PsHD compared to sample representation and pair-wise similarity distillation methods theoretically and experimentally. Finally, experimental results on three widely used datasets validate that the new PsHD outperforms state-of-the-art with 3.9% improvements on average, and also achieves 1.5% improvements while reducing 60% memory buffer size, highlighting the potential of utilizing unlabeled data in SSCL. Our code is available: https://github.com/fanyan0411/PsHD. Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu |
NeurIPS | 2 |
| 2024 | What Matters in Graph Class Incremental Learning? An Information Preservation PerspectiveabstractGraph class incremental learning (GCIL) requires the model to classify emerging nodes of new classes while remembering old classes. Existing methods are designed to preserve effective information of old models or graph data to alleviate forgetting, but there is no clear theoretical understanding of what matters in information preservation. In this paper, we consider that present practice suffers from high semantic and structural shifts assessed by two devised shift metrics. We provide insights into information preservation in GCIL and find that maintaining graph information can preserve information of old models in theory to calibrate node semantic and graph structure shifts. We correspond graph information into low-frequency local-global information and high-frequency information in spatial domain. Based on the analysis, we propose a framework, Graph Spatial Information Preservation (GSIP). Specifically, for low-frequency information preservation, the old node representations obtained by inputting replayed nodes into the old model are aligned with the outputs of the node and its neighbors in the new model, and then old and new outputs are globally matched after pooling. For high-frequency information preservation, the new node representations are encouraged to imitate the near-neighbor pair similarity of old node representations. GSIP achieves a 10\% increase in terms of the forgetting metric compared to prior methods on large-scale datasets. Our framework can also seamlessly integrate existing replay designs. The code is available through https://github.com/Jillian555/GSIP. Yu Wang 0106, Pengfei Zhu 0001, Wanyu Lin, Qinghua Hu |
NeurIPS | 2 |
| 2024 | Integrated Heterogeneous Graph Attention Network for Incomplete Multi-modal Clustering
Yu Wang 0106, Xinjie Yao, Pengfei Zhu 0001, Qinghua Hu |
Int. J. Comput. Vis. | 1 |
| 2024 | Improved generative adversarial network with deep metric learning for missing data imputation
Mohammed Al-taezi, Yu Wang 0106, Pengfei Zhu 0001, Qinghua Hu, Abdulrahman Al-Badwi |
Neurocomputing | 2 |
| 2024 | R2-trans: Fine-grained visual categorization with redundancy reduction
Shuo Ye, Shujian Yu, Yu Wang 0106, Xinge You |
Image Vis. Comput. | 3 |
| 2024 | Few-Shot Learning With Multi-Granularity Knowledge Fusion and Decision-MakingabstractFew-shot learning (FSL) is a challenging task in classifying new classes from few labelled examples. Many existing models embed class structural knowledge as prior knowledge to enhance FSL against data scarcity. However, they fall short of connecting the class structural knowledge with the limited visual information which plays a decisive role in FSL model performance. In this paper, we propose a unified FSL framework with multi-granularity knowledge fusion and decision-making (MGKFD) to overcome the limitation. We aim to simultaneously explore the visual information and structural knowledge, working in a mutual way to enhance FSL. On the one hand, we strongly connect global and local visual information with multigranularity class knowledge to explore intra-image and inter-class relationships, generating specific multi-granularity class representations with limited images. On the other hand, a weight fusion strategy is introduced to integrate multi-granularity knowledge and visual information to make the classification decision of FSL. It enables models to learn more effectively from limited labelled examples and allows generalization to new classes. Moreover, considering varying erroneous predictions, a hierarchical loss is established by structural knowledge to minimize the classification loss, where greater degree of misclassification is penalized more. Experimental results on three benchmark datasets show the advantages of MGKFD over several advanced models. Yuling Su, Hong Zhao 0002, Yifeng Zheng 0004, Yu Wang 0106 |
IEEE Trans. Big Data | 4 |
| 2024 | The Image Data and Backbone in Weakly Supervised Fine-Grained Visual Categorization: A Revisit and Further ThinkingabstractWeakly-supervised fine-grained visual categorization (FGVC) aims to achieve subclass classification within the same large class using only label information. Compared to general images, fine-grained images have similar appearances and features, and are often affected by disturbances such as viewpoint, lighting, and occlusion during data collection, resulting in significant intra-class variance and small inter-class variance. To achieve FGVC, carefully designed models are often needed to explore the locally discriminative regions of the image. This paper revisits high-quality FGVC publications based on deep learning and analyzes from two new perspective: fine-grained image data and backbone. We address two ignored but interesting problems in FGVC. First, we argue that the reasons for exacerbating intra-class variance are not the same in data of animal, plant, and commodity types, and it is necessary to consider the effects of posture, covariate shift, and structural changes. Additionally, the “soft boundary” between subclasses intensifies the difficulty of classification. Second, we highlight that convolutional networks and self-attention networks have different receptive fields and shape biases, leading to performance differences when processing different types of fine-grained data. Overall, our analysis provides new insights into recent advances, challenges, and future directions for FGVC based on deep learning, which can help researchers develop more effective models for FGVC. Shuo Ye, Yu Wang 0106, Qinmu Peng, Xinge You, C. L. Philip Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Multiview Deep Subspace Clustering NetworksabstractMultiview subspace clustering aims to discover the inherent structure of data by fusing multiple views of complementary information. Most existing methods first extract multiple types of handcrafted features and then learn a joint affinity matrix for clustering. The disadvantage of this approach lies in two aspects: 1) multiview relations are not embedded into feature learning and 2) the end-to-end learning manner of deep learning is not suitable for multiview clustering. Even when deep features have been extracted, it is a nontrivial problem to choose a proper backbone for clustering on different datasets. To address these issues, we propose the multiview deep subspace clustering networks (MvDSCNs), which learns a multiview self-representation matrix in an end-to-end manner. The MvDSCN consists of two subnetworks, i.e., a diversity network (Dnet) and a universality network (Unet). A latent space is built using deep convolutional autoencoders, and a self-representation matrix is learned in the latent space using a fully connected layer. Dnet learns view-specific self-representation matrices, whereas Unet learns a common self-representation matrix for all views. To exploit the complementarity of multiview representations, the Hilbert-Schmidt independence criterion (HSIC) is introduced as a diversity regularizer that captures the nonlinear, high-order interview relations. Because different views share the same label space, the self-representation matrices of each view are aligned to the common one by universality regularization. The MvDSCN also unifies multiple backbones to boost clustering performance and avoid the need for model selection. Experiments demonstrate the superiority of the MvDSCN. Pengfei Zhu 0001, Xinjie Yao, Yu Wang 0106, Binyuan Hui, Dawei Du, Qinghua Hu |
IEEE Trans. Cybern. | 3 |
| 2024 | Discriminative Suprasphere Embedding for Fine-Grained Visual CategorizationabstractDespite the great success of the existing work in fine-grained visual categorization (FGVC), there are still several unsolved challenges, e.g., poor interpretation and vagueness contribution. To circumvent this drawback, motivated by the hypersphere embedding method, we propose a discriminative suprasphere embedding (DSE) framework, which can provide intuitive geometric interpretation and effectively extract discriminative features. Specifically, DSE consists of three modules. The first module is a suprasphere embedding (SE) block, which learns discriminative information by emphasizing weight and phase. The second module is a phase activation map (PAM) used to analyze the contribution of local descriptors to the suprasphere feature representation, which uniformly highlights the object region and exhibits remarkable object localization capability. The last module is a class contribution map (CCM), which quantitatively analyzes the network classification decision and provides insight into the domain knowledge about classified objects. Comprehensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed method in comparison with state-of-the-art methods. Shuo Ye, Qinmu Peng, Wenju Sun, Jiamiao Xu, Yu Wang 0106, Xinge You, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Learning Dynamic Compact Memory Embedding for Deformable Visual Object TrackingabstractRecently, template-based trackers have become the leading tracking algorithms with promising performance in terms of efficiency and accuracy. However, the correlation operation between query feature and the given template only achieves accurate target localization, but is prone to state estimation error, especially when the target suffers from severe deformation. To address this issue, segmentation-based trackers are proposed that use per-pixel matching to improve the tracking performance of deformable objects effectively. However, most of the existing trackers only match with the target features of the initial frame, thereby lacking the discrimination for handling a variety of challenging factors, e.g., similar distractors, background clutter, and appearance change. To this end, we propose a dynamic compact memory embedding technique to enhance the discrimination of the segmentation-based visual tracking method that can well tell the target from the background. Specifically, we initialize a memory embedding with the target features in the first frame. During the tracking process, the current target features that have certain correlation with the existing memory are updated to the memory embedding online. To further improve the tracking accuracy for deformable objects, we use a weighted point-to-global matching strategy to measure the correlation between the pixelwise query feature and the whole template, so as to capture more detailed deformation information. Extensive evaluations on six challenging tracking benchmarks including VOT2016, VOT2018, VOT2019, GOT-10K, TrackingNet, and LaSOT demonstrate the superiority of our method over recent remarkable trackers. Besides, our tracker outperforms the excellent segmentation-based trackers, i.e., D3S and SiamMask on the DAVIS2017 benchmark. The code is available at https://github.com/peace-love243/CMEDFL. Pengfei Zhu 0001, Kaihua Zhang 0001, Yu Wang 0106, Tianzhu Zhang 0001, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Coping with change: Learning invariant and minimum sufficient representations for fine-grained visual categorization
Shuo Ye, Shujian Yu, Wenjin Hou, Yu Wang 0106, Xinge You |
Comput. Vis. Image Underst. | 4 |
| 2023 | Class-Specific Semantic Reconstruction for Open Set RecognitionabstractOpen set recognition enables deep neural networks (DNNs) to identify samples of unknown classes, while maintaining high classification accuracy on samples of known classes. Existing methods based on auto-encoder (AE) and prototype learning show great potential in handling this challenging task. In this study, we propose a novel method, called Class-Specific Semantic Reconstruction (CSSR), that integrates the power of AE and prototype learning. Specifically, CSSR replaces prototype points with manifolds represented by class-specific AEs. Unlike conventional prototype-based methods, CSSR models each known class on an individual AE manifold, and measures class belongingness through AE's reconstruction error. Class-specific AEs are plugged into the top of the DNN backbone and reconstruct the semantic representations learned by the DNN instead of the raw image. Through end-to-end learning, the DNN and the AEs boost each other to learn both discriminative and representative information. The results of experiments conducted on multiple datasets show that the proposed method achieves outstanding performance in both close and open set recognition and is sufficiently simple and flexible to incorporate into existing frameworks. Hongzhi Huang, Yu Wang 0106, Qinghua Hu, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | OpenMix+: Revisiting Data Augmentation for Open Set RecognitionabstractOpen set recognition requires models to recognize samples of known classes learned in the training set while reject unknowns not learned. Compared with the structural risk minimization theory for closed-set problems, structural risk in open set tasks remains rarely explored. In this paper, we point out that balancing between structural risk and open space risk is crucial for open set recognition, and re-formalize it as open set structural risk. This brings a new view towards the general relationship between closed set recognition and open set recognition against the common intuition, which argues that a good closed set classifier always benefits for open set recognition. Specifically, we theoretically and experimentally show that recent mix-based data augmentation methods are aggressive closed set regularization methods, which reduce structural risk at cost of sacrificing open space risk. Besides, we show that existing negative data augmentation designed for open space risk reduction also ignore the trade-off problem between structural risk and open space risk, which limits their performance. We propose an efficient negative data augmentation strategy named self-mix and a corresponding method named OpenMix. OpenMix generates high-quality negative samples by mixing samples themselves, which can take care of both risks simultaneously. When combining OpenMix with conservative closed set regularization methods to form OpenMix+, models can achieve lower open set structural risk. Extensive experiments validate the superiority of OpenMix and OpenMix+ in terms of both effectiveness and universality. Guosong Jiang, Pengfei Zhu 0001, Yu Wang 0106, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Multi-Granularity Regularized Re-Balancing for Class Incremental LearningabstractDeep learning models suffer from catastrophic forgetting when learning new tasks incrementally. Incremental learning has been proposed to retain the knowledge of old classes while learning to identify new classes. A typical approach is to use a few exemplars to avoid forgetting old knowledge. In such a scenario, data imbalance between old and new classes is a key issue that leads to performance degradation of the model. Several strategies have been designed to rectify the bias towards the new classes due to data imbalance. However, they heavily rely on the assumptions of the bias relation between old and new classes. Therefore, they are not suitable for complex real-world applications. In this study, we propose an assumption-agnostic method, Multi-Granularity Regularized re-Balancing (MGRB), to address this problem. Re-balancing methods are used to alleviate the influence of data imbalance; however, we empirically discover that they would under-fit new classes. To this end, we further design a novel multi-granularity regularization term that enables the model to consider the correlations of classes in addition to re-balancing the data. A class hierarchy is first constructed by ontology or grouping semantically or visually similar classes. The multi-granularity regularization then transforms the one-hot label vector into a continuous label distribution, which reflects the relations between the target class and other classes based on the constructed class hierarchy. Thus, the model can learn the inter-class relational information, which helps enhance the learning of both old and new classes. Experimental results on both public datasets and a real-world fault diagnosis dataset verify the effectiveness of the proposed method. Code is available athttps://github.com/lilyht/CIL-MGRB. Huitong Chen, Yu Wang 0106, Qinghua Hu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Robust Multi-Drone Multi-Target Tracking to Resolve Target Occlusion: A BenchmarkabstractMulti-drone multi-target tracking aims at collabo- ratively detecting and tracking targets across multiple drones and associating the identities of objects from different drones, which can overcome the shortcomings of single-drone object tracking. To address the critical challenges of identity association and target occlusion in multi-drone multi-target tracking tasks, we collect an occlusion-aware multi-drone multi-target tracking dataset named MDMT. It contains 88 video sequences with 39,678 frames, including 11,454 different IDs of persons, bicycles, and cars. The MDMT dataset comprises 2,204,620 bounding boxes, of which 543,444 bounding boxes contain target occlusions. We also design a multi-device target association score (MDA) as the evaluation criteria for the ability of cross-view target association in multi-device tracking. Furthermore, we propose a Multi-matching Identity Authentication network (MIA-Net) for the multi-drone multi-target tracking task. The local-global matching algorithm in MIA-Net discovers the topological relationship of targets across drones, efficiently solves the problem of cross-drone association, and effectively complements occluded targets with the advantage of multiple drone view mapping. Extensive experiments on the MDMT dataset validate the effectiveness of our proposed MIA-Net for the task of identity association and multi-object tracking with occlusions. Timing Li, Yu Wang 0106, Qinghua Hu, Pengfei Zhu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Latent Heterogeneous Graph Network for Incomplete Multi-View LearningabstractMulti-view learning has progressed rapidly in recent years. Although many previous studies assume that each instance appears in all views, it is common in real-world applications for instances to be missing from some views, resulting in incomplete multi-view data. To tackle this problem, we propose a novel Latent Heterogeneous Graph Network (LHGN) for incomplete multi-view learning, which aims to use multiple incomplete views as fully as possible in a flexible manner. By learning a unified latent representation, a trade-off between consistency and complementarity among different views is implicitly realized. To explore the complex relationship between samples and latent representations, a neighborhood constraint and a view-existence constraint are proposed, for the first time, to construct a heterogeneous graph. Finally, to avoid any inconsistencies between training and test phase, a transductive learning technique is applied based on graph learning for classification tasks. Extensive experimental results on real-world datasets demonstrate the effectiveness of our model over existing state-of-the-art approaches. Our code is available at:https://github.com/yxjdarren/LHGN_TMM_2022. Pengfei Zhu 0001, Xinjie Yao, Yu Wang 0106, Binyuan Hui, Qinghua Hu |
IEEE Trans. Multim. | 3 |
| 2023 | Coarse-to-Fine: Progressive Knowledge Transfer-Based Multitask Convolutional Neural Network for Intelligent Large-Scale Fault DiagnosisabstractIn modern industry, large-scale fault diagnosis of complex systems is emerging and becoming increasingly important. Most deep learning-based methods perform well on small number of fault diagnosis, but cannot converge to satisfactory results when handling large-scale fault diagnosis because the huge number of fault types will lead to the problems of intra/inter-class distance unbalance and poor local minima in neural networks. To address the above problems, a progressive knowledge transfer-based multitask convolutional neural network (PKT-MCNN) is proposed. First, to construct the coarse-to-fine knowledge structure intelligently, a structure learning algorithm is proposed via clustering fault types in different coarse-grained nodes. Thus, the intra/inter-class distance unbalance problem can be mitigated by spreading similar tasks into different nodes. Then, an MCNN architecture is designed to learn the coarse and fine-grained task simultaneously and extract more general fault information, thereby pushing the algorithm away from poor local minima. Last but not least, a PKT algorithm is proposed, which can not only transfer the coarse-grained knowledge to the fine-grained task and further alleviate the intra/inter-class distance unbalance in feature space, but also regulate different learning stages by adjusting the attention weight to each task progressively. To verify the effectiveness of the proposed method, a dataset of a nuclear power system with 66 fault types was collected and analyzed. The results demonstrate that the proposed method can be a promising tool for large-scale fault diagnosis. Yu Wang 0106, Di Lin 0002, Ping Li 0016, Qinghua Hu, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Collaborative Decision-Reinforced Self-Supervision for Attributed Graph ClusteringabstractAttributed graph clustering aims to partition nodes of a graph structure into different groups. Recent works usually use variational graph autoencoder (VGAE) to make the node representations obey a specific distribution. Although they have shown promising results, how to introduce supervised information to guide the representation learning of graph nodes and improve clustering performance is still an open problem. In this article, we propose a Collaborative Decision-Reinforced Self-Supervision (CDRS) method to solve the problem, in which a pseudo node classification task collaborates with the clustering task to enhance the representation learning of graph nodes. First, a transformation module is used to enable end-to-end training of existing methods based on VGAE. Second, the pseudo node classification task is introduced into the network through multitask learning to make classification decisions for graph nodes. The graph nodes that have consistent decisions on clustering and pseudo node classification are added to a pseudo-label set, which can provide fruitful self-supervision for subsequent training. This pseudo-label set is gradually augmented during training, thus reinforcing the generalization capability of the network. Finally, we investigate different sorting strategies to further improve the quality of the pseudo-label set. Extensive experiments on multiple datasets show that the proposed method achieves outstanding performance compared with state-of-the-art methods. Our code is available at https://github.com/Jillian555/TNNLS_CDRS. Pengfei Zhu 0001, Yu Wang 0106, Bin Xiao 0002, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | RCANet: Row-Column Attention Network for Semantic SegmentationabstractEstablishing high-order interactions among pixels and object parts is one of the most fundamental problems in semantic segmentation. The recent proposals are based on non-local methods which utilize the self-attention mechanism to capture the long-range correlations. However, non-local methods could be very expensive, both theoretically and experimentally. Moreover, non-local methods are typically designed to address spatial correlations rather than feature correlations across channels. In this work, we propose a Row-Column Attention Network (RCANet) to encode globally contextual information. It consists of a row-wise intra-channel attention module and a column-wise intra-channel attention module, followed by a cross-channel interaction module. We conduct experiments on two datasets: Cityscapes and ADE20K. The results show that our method is comparable to the state-of-the-art methods for semantic segmentation. Bingxu Lu, Qinghua Hu, Yu Wang 0106, Guosheng Hu |
ICASSP | 3 |
| 2022 | Learning Self-supervised Low-Rank Network for Single-Stage Weakly and Semi-supervised Semantic Segmentation
Junwen Pan, Pengfei Zhu 0001, Kaihua Zhang 0001, Bing Cao 0002, Yu Wang 0106, Dingwen Zhang, Junwei Han 0001, Qinghua Hu |
Int. J. Comput. Vis. | 5 |
| 2022 | Uncertainty instructed multi-granularity decision for large-scale hierarchical classification
Yu Wang 0106, Qinghua Hu, Hao Chen 0112 |
Inf. Sci. | 1 |
| 2022 | Deep collaborative multi-task network: A human decision process inspired model for hierarchical image classification
Yu Zhou 0015, Xiaoni Li, Yucan Zhou, Yu Wang 0106, Qinghua Hu, Weiping Wang 0005 |
Pattern Recognit. | 4 |
| 2022 | Multi-granularity episodic contrastive learning for few-shot learning
Pengfei Zhu 0001, Yu Wang 0106, Jinglin Zhang 0001 |
Pattern Recognit. | 3 |
| 2022 | Hierarchical Semantic Risk Minimization for Large-Scale ClassificationabstractHierarchical structures of labels usually exist in large-scale classification tasks, where labels can be organized into a tree-shaped structure. The nodes near the root stand for coarser labels, while the nodes close to leaves mean the finer labels. We label unseen samples from the root node to a leaf node, and obtain multigranularity predictions in the hierarchical classification. Sometimes, we cannot obtain a leaf decision due to uncertainty or incomplete information. In this case, we should stop at an internal node, rather than going ahead rashly. However, most existing hierarchical classification models aim at maximizing the percentage of correct predictions, and do not take the risk of misclassifications into account. Such risk is critically important in some real-world applications, and can be measured by the distance between the ground truth and the predicted classes in the class hierarchy. In this work, we utilize the semantic hierarchy to define the classification risk and design an optimization technique to reduce such risk. By defining the conservative risk and the precipitant risk as two competing risk factors, we construct the balanced conservative/precipitant semantic (BCPS) risk matrix across all nodes in the semantic hierarchy with user-defined weights to adjust the tradeoff between two kinds of risks. We then model the classification process on the semantic hierarchy as a sequential decision-making task. We design an algorithm to derive the risk-minimized predictions. There are two modules in this model: 1) multitask hierarchical learning and 2) deep reinforce multigranularity learning. The first one learns classification confidence scores of multiple levels. These scores are then fed into deep reinforced multigranularity learning for obtaining a global risk-minimized prediction with flexible granularity. Experimental results show that the proposed model outperforms state-of-the-art methods on seven large-scale classification datasets with the semantic tree. Yu Wang 0106, Zhou Wang 0001, Qinghua Hu, Yucan Zhou, Honglei Su |
IEEE Trans. Cybern. | 1 |
| 2021 | Get to the Point: Content Classification of Animated Graphics Interchange Formats with Key-Frame AttentionabstractAnimated Graphics Interchange Formats (GIFS) are low-bandwidth short image sequences that can continuously display multiple frames without sound. In this paper, we focus on a new content classification task that is important in real-world applications. A key problem for this task is that some frames in an animated GF are irrelevant to the label, which may drastically reduce the classification performance. To this end, we first collect a new dataset of Web animated GIFS (WGF) that includes some typical samples in which only several key-frames are relevant to the ground truth. Then, an attention-based method is designed to learn to produce importance scores of the frames, and subsequently multi-frame predicted scores are merged to obtain the final prediction. Besides, an additional entropy loss is also used to sharpen the attention results to further emphasize the key-frames. Experimental results on WGF show that the proposed approach significantly outperforms various baseline methods. Yongjuan Ma, Yu Wang 0106, Pengfei Zhu 0001, Junwen Pan |
ICIP | 2 |
| 2021 | PQA-Net: Deep No Reference Point Cloud Quality Assessment via Multi-View ProjectionabstractRecently, 3D point cloud is becoming popular due to its capability to represent the real world for advanced content modality in modern communication systems. In view of its wide applications, especially for immersive communication towards human perception, quality metrics for point clouds are essential. Existing point cloud quality evaluations rely on a full or certain portion of the original point cloud, which severely limits their applications. To overcome this problem, we propose a novel deep learning-based no reference point cloud quality assessment method, namely PQA-Net. Specifically, the PQA-Net consists of a multi-view-based joint feature extraction and fusion (MVFEF) module, a distortion type identification (DTI) module, and a quality vector prediction (QVP) module. The DTI and QVP modules share the feature generated from the MVFEF module. By using the distortion type labels, the DTI and the MVFEF modules are first pre-trained to initialize the network parameters, based on which the whole network is then jointly trained to finally evaluate the point cloud quality. Experimental results on the Waterloo Point Cloud dataset show that PQA-Net achieves better or equivalent performance comparing with the state-of-the-art quality assessment methods. The code of the proposed model will be made publicly available to facilitate reproducible researchhttps://github.com/qdushl/PQA-Net. Qi Liu 0029, Hui Yuan 0001, Honglei Su, Hao Liu 0044, Yu Wang 0106, Huan Yang 0001, Junhui Hou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | A Recursive Regularization Based Feature Selection Framework for Hierarchical ClassificationabstractThe sizes of datasets in terms of the number of samples, features, and classes have dramatically increased in recent years. In particular, there usually exists a hierarchical structure among class labels as hundreds of classes exist in a classification task. We call these tasks hierarchical classification, and hierarchical structures are helpful for dividing a very large task into a collection of relatively small subtasks. Various algorithms have been developed to select informative features for flat classification. However, these algorithms ignore the semantic hyponymy in the directory of hierarchical classes, and select a uniform subset of the features for all classes. In this paper, we propose a new feature selection framework with recursive regularization for hierarchical classification. This framework takes the hierarchical information of the class structure into account. In contrast to flat feature selection, we select different feature subsets for each node in a hierarchical tree structure with recursive regularization. The proposed framework uses parent-child, sibling, and family relationships for hierarchical regularization. By imposing$\ell _{2,1}$-norm regularization to different parts of the hierarchical classes, we can learn a sparse matrix for the feature ranking at each node. Extensive experiments on public datasets demonstrate the effectiveness and efficiency of the proposed algorithms. Hong Zhao 0002, Qinghua Hu, Pengfei Zhu 0001, Yu Wang 0106, Ping Wang 0072 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Deep Fuzzy Tree for Large-Scale Hierarchical Visual ClassificationabstractDeep learning models often use a flat softmax layer to classify samples after feature extraction in visual classification tasks. However, it is hard to make a single decision of finding the true label from massive classes. In this scenario, hierarchical classification is proved to be an effective solution and can be utilized to replace the softmax layer. A key issue of hierarchical classification is to construct a good label structure, which is very significant for classification performance. Several works have been proposed to address the issue, but they have some limitations and are almost designed heuristically. In this article, inspired by fuzzy rough set theory, we propose a deep fuzzy tree model which learns a better tree structure and classifiers for hierarchical classification with theory guarantee. Experimental results show the effectiveness and efficiency of the proposed model in various visual classification datasets. Yu Wang 0106, Qinghua Hu, Pengfei Zhu 0001, Linhao Li, Bingxu Lu, Jonathan M. Garibaldi, Xianling Li |
IEEE Trans. Fuzzy Syst. | 1 |
| 2018 | Monotonicity Extraction for Monotonic Bayesian Networks Parameter Learning
Jingzhuo Yang, Yu Wang 0106, Qinghua Hu |
ICONIP (3) | 2 |
| 2018 | Monotonicity Induced Parameter Learning for Bayesian Networks with Limited DataabstractParameter learning of Bayesian networks (BNs) is a challenging task as it depends heavily on a large number of reliable training data. Unfortunately, it is often difficult to obtain sufficient samples in many real-world applications. Fortunately, monotonicity relationship among variables widely exists in many practical tasks, and has been proven to be effective in learning the parameters of BN with limited data. Most researches utilize monotonicity relationship provided manually by domain experts, but it is difficult and costly to obtain all the prior knowledge of monotonicity accurately if the structure of BN is quite complex. In this paper, we propose a data-dependent method to learn the parameters of BN with limited data. Firstly, Spearman rank correlation coefficient (RHO) is leverages to detect the monotonicity relationship between the network nodes. Secondly, the monotonicity relationship is transformed into a set of monotonicity constraints for the network parameters, and then integrated into the log-likehood function as a penalty item (RHO-PML). Finally, the parameters of BN are obtained by the gradient descent method. Moreover, to reinforce the impact of the monotonicity relationship, bidirectional monotonicity constraints are introduced into RHO-PML as RHO-BPML. Experiments on various datasets show the effectiveness of the proposed RHO-PML and RHO-BPML algorithms with limited data. Jingzhuo Yang, Yu Wang 0106, Shenglei Pei, Qinghua Hu |
IJCNN | 2 |
| 2018 | Deep super-class learning for long-tail distributed image classification
Yucan Zhou, Qinghua Hu, Yu Wang 0106 |
Pattern Recognit. | 3 |
| 2017 | Local Bayes Risk Minimization Based Stopping Strategy for Hierarchical ClassificationabstractIn large-scale data classification tasks, it is becoming more and more challenging in finding a true class from a huge amount of candidate categories. Fortunately, a hierarchical structure usually exists in these massive categories. The task of utilizing this structure for effective classification is called hierarchical classification. It usually follows a top-down fashion which predicts a sample from the root node with a coarse-grained category to a leaf node with a fine-grained category. However, misclassification is inevitable if the information is insufficient or large uncertainty exists in the prediction process. In this scenario, we can design a stopping strategy to stop the sample at an internal node with a coarser category, instead of predicting a wrong leaf node. Several studies address the problem by improving performance in terms of hierarchical accuracy and informative prediction. However, all of these researches ignore an important issue: when predicting a sample at the current node, the error is inclined to occur if large uncertainty exists in the next lower level children nodes. In this paper, we integrate this uncertainty into a risk problem: when predicting a sample at a decision node, it will take precipitance risk in predicting the sample to a children node in the next lower level on one hand, and take conservative risk in stopping at the current node on the other. We address the risk problem by designing a Local Bayes Risk Minimization (LBRM) framework, which divides the prediction process into recursively deciding to stop or to go down at each decision node by balancing these two risks in a top-down fashion. Rather than setting a global loss function in the traditional Bayes risk framework, we replace it with different uncertainty in the two risks for each decision node. The uncertainty on the precipitance risk and the conservative risk are measured by information entropy on children nodes and information gain from the current node to children nodes, respectively. We propose a Weighted Tree Induced Error (WTIE) to obtain the predictions of minimum risk with different emphasis on the two risks. Experimental results on various datasets show the effectiveness of the proposed LBRM algorithm. Yu Wang 0106, Qinghua Hu, Yucan Zhou, Hong Zhao 0002, Jiye Liang |
ICDM | 1 |