Yan Zhang 0053

dblp:04/3348-53 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
11since 2021 · last 2026
0009-0006-6418-5296ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction
abstract
Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models through robot interaction. However, SAGCI suffers from prior-dependency (requiring initialization), neglects kinematic/dynamic constraints, and generates non-watertight meshes. To overcome these limitations, we propose SIAM, a novel framework for efficient and generalizable Single-Interaction Articulated Modeling. Given an initial point cloud, SIAM first enables minimal robot interaction to trigger object motion. It then precisely segments parts by analyzing point cloud differences pre- and post-interaction. For joint parameter estimation, we introduce an optimization incorporating novel kinematic energy constraints, enhancing physical consistency. Finally, we reconstruct a high-quality, topologically watertight mesh by learning 3D Gaussian Primitives from multi-view RGB-D observations under deformation. Extensive experiments on the PartNet-Mobility benchmark demonstrate state-of-the-art articulation modeling performance. Successful real-world deployment with an xArm robot further validates the framework's practicality and transferability. SIAM achieves accurate, prior-free modeling with significantly reduced interaction cost.
Li Zhang 0104, Yan Zhang 0053, Anran Huang, Liu Liu 0012, Dan Guo 0001
AAAI4
2026 Multi-Class Object Counting Network With Adaptive Class Alignment Loss for Remote Sensing Images
abstract
Accurately estimating the number of objects across different categories is crucial in remote sensing images for applications ranging from urban planning and environmental monitoring to disaster management. Although there are a few studies about multi-class object counting for remote sensing images, these methods may suffer from the issues of class imbalance, inter-class interference, or inability to capture fine-grained features. Thus, we propose a novel network called Multi-class Object Counting Network with Adaptive Class Alignment loss (ACA-MOCN) for multi-class object counting in remote sensing images. ACA-MOCN newly designs a Adaptive Class Alignment (ACA) loss function that can effectively balance the losses across multi-class objects and reduce interference between classes. In addition, ACA-MOCN also designs a Filtering Feature Pyramid Network (FFPN) structure to facilitate an effective feature fusion mechanism, enabling the integration of granular and detailed information. Experimental results on the two public datasets demonstrate that ACA-MOCN is superior when handling large-scale multi-class object counting tasks in remote sensing images.
Li Zhang 0004, Hao-Yuan Ma, Xiang-Yi Wei, Yan Zhang 0053
IEEE Signal Process. Lett.5
2025 Discrimination-Enhanced Hierarchical Network with Structure Optimization for Cross-Subject EEG Emotion Recognition
abstract
At present, EEG-based emotion recognition has advanced significantly. However, models' discrimination ability across subjects remains limited because of ignoring the electrode structure and the cross-subject differences in EEG signals. To address this, we propose a discrimination-enhanced hierarchical network with structure optimization (SO-DHNet), a discrimination-enhanced hierarchical network with structure optimization for cross-subject EEG emotion recognition. It includes a hierarchical data generation (HDG) module that organizes EEG electrodes into five neuroanatomical regions, optimizes their structure, and generates 4D hierarchical data. Additionally, we introduce a random pairing training (RPT) strategy with a pairwise-similarity binary cross-entropy (PS-BCE) loss to maximize inter-class diversity and minimize intra-class variability. Experiments on SEED confirm SO-DHNet's effectiveness in cross-subject emotion recognition.
Yilin Wang 0003, Li Zhang 0004, Yan Zhang 0053
BIBM3
2025 Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
abstract
Micro-Action Recognition (MAR) aims to classify subtle human actions in video. However, annotating MAR datasets is particularly challenging due to the subtlety of actions. To this end, we introduce the setting of Semi-Supervised MAR (SSMAR), where only a part of samples are labeled. We first evaluate traditional Semi-Supervised Learning (SSL) methods to SSMAR and find that these methods tend to overfit on inaccurate pseudo-labels, leading to error accumulation and degraded performance. This issue primarily arises from the common practice of directly using the predictions of classifier as pseudo-labels to train the model. To solve this issue, we propose a novel framework, called Asynchronous Pseudo Labeling and Training (APLT), which explicitly separates the pseudo-labeling process from model training. Specifically, we introduce a semi-supervised clustering method during the offline pseudo-labeling phase to generate more accurate pseudo-labels. Moreover, a self-adaptive thresholding strategy is proposed to dynamically filter noisy labels of different classes. We then build a memory-based prototype classifier based on the filtered pseudo-labels, which is fixed and used to guide the subsequent model training phase. By alternating the two pseudo-labeling and model training phases in an asynchronous manner, the model can not only be learned with more accurate pseudo-labels but also avoid the overfitting issue. Experiments on three MAR datasets show that our APLT largely outperforms state-of-the-art SSL methods. For instance, APLT improves accuracy by 14.5% over FixMatch on the MA-12 dataset when using only 50% labeled data. Code is available at https://github.com/zy-hfut/APLT
Yan Zhang 0053, Lechao Cheng, Yaxiong Wang, Zhun Zhong, Meng Wang 0001
IJCAI1
2025 MAGIC: Noise Mitigation and Knowledge Alignment for Knowledge Graph-Based Multi-modal Recommendation
abstract
Multi-modal recommender systems (MMRSs) have demonstrated significant potential in mitigating data sparsity and cold start problems by leveraging diverse multi-modal data, such as text and images. To further improve the recommendation accuracy, some MMRSs have integrated knowledge graphs (KGs) to enrich the graph structure with meaningful relationships between entities, giving rise to the task of KG-based MMRSs. Despite the promising results achieved by existing studies on this task, (i) they overlook the substantial noise introduced within the auxiliary information (i.e., both KGs and multi-modal data), and (ii) most of them struggle to effectively align the knowledge from history user-item interactions and auxiliary information. To tackle these limitations, we propose a novel model entitled MAGIC (noise Mitigation and knowledge Aignment for knowledge Graph-based multI-modal reCommendation). Specifically, to tackle the limitation (i), we design a noise-aware heterogeneous aggregation layer in the KG-based modal enhancement module. To address the limitation (ii), MAGIC introduces adversarial learning in the CF-based adversarial learning module, and exploits contrastive learning in the fusion and prediction module. The experiments conducted on two extended real-world datasets from different domains demonstrate the superiority of MAGIC over state-of-the-art baselines.
Yan Zhang 0053, Li Zhang 0004, Xi Chen 0121, Lei Zhao 0001
ICMR2
2024 Benchmarking Micro-Action Recognition: Dataset, Methods, and Applications
abstract
Micro-action is an imperceptible non-verbal behaviour characterised by low-intensity movement. It offers insights into the feelings and intentions of individuals and is important for human-oriented applications such as emotion recognition and psychological assessment. However, the identification, differentiation, and understanding of micro-actions pose challenges due to the imperceptible and inaccessible nature of these subtle human behaviors in everyday life. In this study, we innovatively collect a new micro-action dataset designated as Micro-action-52 (MA-52), and propose a benchmark named micro-action network (MANet) for micro-action recognition (MAR) task. Uniquely, MA-52 provides the whole-body perspective including gestures, upper- and lower-limb movements, attempting to reveal comprehensive micro-action cues. In detail, MA-52 contains 52 micro-action categories along with seven body part labels, and encompasses a full array of realistic and natural micro-actions, accounting for 205 participants and 22,422 video instances collated from the psychological interviews. Based on the proposed dataset, we assess MANet and other nine prevalent action recognition methods. MANet incorporates squeeze-and-excitation (SE) and temporal shift module (TSM) into the ResNet architecture for modeling the spatiotemporal characteristics of micro-actions. Then a joint-embedding loss is designed for semantic matching between video and action labels; the loss is used to better distinguish between visually similar yet distinct micro-action categories. The extended application in emotion recognition has demonstrated one of the important values of our proposed dataset and method. In the future, further exploration of human behaviour, emotion, and psychological assessment will be conducted in depth. The dataset and source code are released at https://github.com/VUT-HFUT/Micro-Action.
Dan Guo 0001, Kun Li 0008, Bin Hu 0001, Yan Zhang 0053, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Spatiotemporal contrastive modeling for video moment retrieval
Kun Li 0008, Yan Zhang 0053, Dan Guo 0001, Meng Wang 0001
World Wide Web (WWW)4
2021 Partial-Label and Structure-constrained Deep Coupled Factorization Network
abstract
In this paper, we technically propose an enriched prior guided framework, called Dual-constrained Deep Semi-Supervised Coupled Factorization Network (DS2CF-Net), for discovering hierarchical coupled data representation. To extract hidden deep features, DS2CF-Net is formulated as a partial-label and geometrical structure-constrained framework. Specifically, DS2CF-Net designs a deep factorization architecture using multilayers of linear transformations, which can coupled update both the basis vectors and new representations in each layer. To enable learned deep representations and coefficients to be discriminative, we also consider enriching the supervised prior by joint deep coefficients-based label prediction and then incorporate the enriched prior information as additional label and structure constraints. The label constraint can enable the intra-class samples to have same coordinate in feature space, and the structure constraint forces the coefficients in each layer to be block-diagonal so that the enriched prior using the self-expressive label propagation are more accurate. Our network also integrates the adaptive dual-graph learning to retain the local structures of both data and feature manifolds in each layer. Extensive experiments on image datasets demonstrate the effectiveness of DS2CF-Net for representation learning and clustering.
Yan Zhang 0053, Zhao Zhang 0001, Yang Wang 0023, Zheng Zhang 0006, Li Zhang 0004, Shuicheng Yan, Meng Wang 0001
AAAI1
2021 Dual-Constrained Deep Semi-Supervised Coupled Factorization Network with Enriched Prior
Yan Zhang 0053, Zhao Zhang 0001, Yang Wang 0023, Zheng Zhang 0006, Li Zhang 0004, Shuicheng Yan, Meng Wang 0001
Int. J. Comput. Vis.1
2021 A Survey on Concept Factorization: From Shallow to Deep Representation Learning
Zhao Zhang 0001, Yan Zhang 0053, Mingliang Xu 0001, Li Zhang 0004, Yi Yang 0001, Shuicheng Yan
Inf. Process. Manag.2
2021 Flexible Auto-Weighted Local-Coordinate Concept Factorization: A Robust Framework for Unsupervised Clustering
abstract
Concept Factorization (CF) and its variants may produce inaccurate representation and clustering results due to the sensitivity to noise, hard constraint on the reconstruction error, and pre-obtained approximate similarities. To improve the representation ability, a novel unsupervised Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF) framework is proposed for clustering high-dimensional data. Specifically, RFA-LCF integrates the robust flexible CF by clean data space recovery, robust sparse local-coordinate coding, and adaptive weighting into a unified model. RFA-LCF improves the representations by enhancing the robustness of CF to noise and errors, providing a flexible constraint on the reconstruction error and optimizing the locality jointly. For robust learning, RFA-LCF clearly learns a sparse projection to recover the underlying clean data space, and then the flexible CF is performed in the projected feature space. RFA-LCF also uses a L2,1-norm based flexible residue to encode the mismatch between the recovered data and its reconstruction, and uses the robust sparse local-coordinate coding to represent data using a few nearby basis concepts. For auto-weighting, RFA-LCF jointly preserves the manifold structures in the basis concept space and new coordinate space in an adaptive manner by minimizing the reconstruction errors on clean data, anchor points and coordinates. By updating the local-coordinate preserving data, basis concepts and new coordinates alternately, the representation abilities can be potentially improved. Extensive results on public databases show that RFA-LCF delivers the state-of-the-art clustering results compared with other related methods.
Zhao Zhang 0001, Yan Zhang 0053, Sheng Li 0001, Guangcan Liu, Dan Zeng 0001, Shuicheng Yan, Meng Wang 0001
IEEE Trans. Knowl. Data Eng.2
2020 Deep Self-representative Concept Factorization Network for Representation Learning
abstract
In this paper, we technically propose a novel framework called Deep Self-representative Concept Factorization Network (DSCF-Net), for clustering deep features. To improve the representation and clustering abilities, DSCF-Net explicitly considers discovering hidden deep semantic features, enhancing the robustness properties of the deep factorization to noise and preserving the local manifold structures of deep features. Specifically, DSCF-Net integrates the robust deep concept factorization, deep self-expressive representation and adaptive locality preserving feature learning into a unified framework. To discover hidden deep representations, DSCF-Net designs a hierarchical factorization architecture using multiple layers of linear transformations, where the hierarchical representation is performed by formulating the problem as optimizing the basis concepts in each layer to improve the representation indirectly. DSCF-Net also improves robustness by subspace recovery for sparse error correction firstly and then performs deep factorization in the recovered visual subspace. To obtain localitypreserving representations, we also present an adaptive deep self-representative weighting strategy by using the coefficient matrix as adaptive weights to keep the locality of representations. Extensive results show that DSCF-Net delivers state-of-the-art performance on several public databases.
Yan Zhang 0053, Zhao Zhang 0001, Zheng Zhang 0006, Ming-Bo Zhao, Li Zhang 0004, Zhengjun Zha, Meng Wang 0001
SDM1
2020 Joint Label Prediction Based Semi-Supervised Adaptive Concept Factorization for Robust Data Representation
abstract
Constrained Concept Factorization (CCF) yields the enhanced representation ability over CF by incorporating label information as additional constraints, but it cannot classify and group unlabeled data appropriately. Minimizing the difference between the original data and its reconstruction directly can enable CCF to model a small noisy perturbation, but is not robust to gross sparse errors. Besides, CCF cannot preserve the manifold structures in new representation space explicitly, especially in an adaptive manner. In this paper, we propose a joint label prediction based Robust Semi-Supervised Adaptive Concept Factorization (RS2ACF) framework. To obtain robust representation, RS2ACF relaxes the factorization to make it simultaneously stable to small entrywise noise and robust to sparse errors. To enrich prior knowledge to enhance the discrimination, RS2ACF clearly uses class information of labeled data and more importantly propagates it to unlabeled data by jointly learning an explicit label indicator for unlabeled data. By the label indicator, RS2ACF can ensure the unlabeled data of the same predicted label to be mapped into the same class in feature space. Besides, RS2ACF incorporates the joint neighborhood reconstruction error over the new representations and predicted labels of both labeled and unlabeled data, so the manifold structures can be preserved explicitly and adaptively in the representation space and label space at the same time. Owing to the adaptive manner, the tricky process of determining the neighborhood size or kernel width can be avoided. Extensive results on public databases verify that our RS2ACF can deliver state-of-the-art data representation, compared with other related methods.
Zhao Zhang 0001, Yan Zhang 0053, Guangcan Liu, Jinhui Tang 0001, Shuicheng Yan, Meng Wang 0001
IEEE Trans. Knowl. Data Eng.2
2019 Robust Unsupervised Flexible Auto-weighted Local-coordinate Concept Factorization for Image Clustering
abstract
We investigate the high-dimensional data clustering problem by proposing a novel and unsupervised representation learning model called Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF). RFA-LCF integrates the robust flexible CF, robust sparse local-coordinate coding and the adaptive reconstruction weighting learning into a unified model. The adaptive weighting is driven by including thejoint manifold preserving constraints on the recovered clean data, basis concepts and new representation. Specifically, our RFA-LCF uses a L2,1-norm based flexible residue to encode the mismatch between clean data and its reconstruction, and also applies the robust adaptive sparse local-coordinate coding to represent the data using a few nearby basis concepts, which can make the factorization more accurate and robust to noise. The robust flexible factorization is also performed in the recovered clean data space for enhancing representations. RFA-LCF also considers preserving the local manifold structures of clean data space, basis concept space and the new coordinate space jointly in an adaptive manner way. Extensive comparisons show that RFA-LCF can deliver enhanced clustering results.
Zhao Zhang 0001, Yan Zhang 0053, Sheng Li 0001, Guangcan Liu, Meng Wang 0001, Shuicheng Yan
ICASSP2
2019 Unsupervised Nonnegative Adaptive Feature Extraction for Data Representation
abstract
In this paper, we propose a novel unsupervised Nonnegative Adaptive Feature Extraction (NAFE) algorithm for data representation and classification. The formulation of NAFE integrates the sparsity constrained nonnegative matrix factorization (NMF), representation learning, and adaptive reconstruction weight learning into a unified model. Specifically, NAFE performs feature and weight learning over the new robust representations of NMF for more accurate measure and representation. For nonnegative adaptive feature extraction, our NAFE first utilizes the sparsity constrained NMF to obtain the new and robust representations of the original data. To preserve the manifold structures of the learnt new representations, we also incorporate a neighborhood reconstruction error over the weight matrix for joint minimization. Note that to further improve the representation power, the weights are jointly shared in the new low-dimensional nonnegative representation space, low-dimensional nonlinear manifold space, and low-dimensional projective subspace, i.e., local neighborhood information is clearly preserved in different feature spaces so that informative representations and features can be jointly obtained. To enable NAFE to extract features from new data, we also include a feature approximation error by a linear projection so that the learnt extractor can obtain features from new data efficiently. Extensive simulations show that our formulation can deliver state-of-the-art results on several public databases for feature extraction and classification, compared with several related methods.
Yan Zhang 0053, Zhao Zhang 0001, Sheng Li 0001, Jie Qin 0004, Guangcan Liu, Meng Wang 0001, Shuicheng Yan
IEEE Trans. Knowl. Data Eng.1
2018 Semi-supervised local multi-manifold Isomap by linear embedding for feature extraction
Yan Zhang 0053, Zhao Zhang 0001, Jie Qin 0004, Li Zhang 0004, Bing Li 0007, Fanzhang Li
Pattern Recognit.1
2017 Discriminative sparse flexible manifold embedding with novel graph for robust visual representation and label propagation
Zhao Zhang 0001, Yan Zhang 0053, Fanzhang Li, Ming-Bo Zhao, Li Zhang 0004, Shuicheng Yan
Pattern Recognit.2
2016 Semi-supervised Classification by Nuclear-Norm Based Transductive Label Propagation
Lei Jia 0002, Zhao Zhang 0001, Yan Zhang 0053
ICONIP (3)3
2016 Robust Soft Semi-supervised Discriminant Projection for Feature Learning
Zhao Zhang 0001, Yan Zhang 0053
ICONIP (2)3
2016 Robust L1-norm matrixed locality preserving projection for discriminative subspace learning
abstract
L1-norm maximization based Discriminant Locality Preserving Projection (DLPP-L1) is shown to be effective and robust to the outliers in given data, but DLPP-L1 is based on the vector space, so it has to convert those 2D matrices into high-dimensional 1D vectorized representations when handing images. But such transformation usually destroys the topology structures of images pixels, which can decrease performance. We therefore propose to extend DLPP-L1 to the 2D matrix space. A two-dimensional DLPP-L1, termed 2D-DLPP-L1, is technically proposed for image feature extraction. Compared with DLPP-L1 for representation, our proposed 2D-DLPP-L1 can effectively preserve the topology structures among image pixels in addition to inheriting the robustness property against noise and outliers. Extensive simulations on real-world image datasets show that our 2D-DLPP-L1 can deliver enhanced performance over other state-of-the-arts for recognition.
Zhao Zhang 0001, Yan Zhang 0053, Fanzhang Li
IJCNN3
2016 Joint nuclear-norm nonlinear Manifold Learning and robust Classification by linear embedding
abstract
We propose a Joint nuclear norm based nonlinear Manifold Learning through linear embedding with Classification, called JMLC. By including a feature approximation error into the existing nonlinear manifold learning framework to correlate manifold features with embedded features by a linear projection, the learnt projection can handle the outside points efficiently by embedding. Besides, to encode the neighborhood reconstruction error in manifold learning part, we apply a more reliable nuclear norm based distance metric, since nuclear norm is proved to be more reliable than both L1-norm and Frobenius norm. To make learnt nonlinear features be optimal for classification so that the accuracy can be enhanced, we also minimize a robust L2,1-norm based regressive classification error over the embedded manifold features, which ensures the classification process to be robust to noise and outliers in data. Based on performing joint manifold learning and classification alternately, our JMLC can obtain a low-dimensional embedding, a linear projection and a multi-class classifier simultaneously. Extensive results demonstrate the validity of our proposed algorithm for feature extraction and robust classification, compared with other related models.
Yan Zhang 0053, Zhao Zhang 0001, Lei Jia 0002, Fanzhang Li
IJCNN1