Zan Gao 0002

dblp:02/2658-2 · DBLP profile ↗
← Back
30ranked-venue papers
13as first author
11since 2021 · last 2026
0000-0001-6970-491XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Separating Domain-Private Classes for Universal Unsupervised Cross-Domain 3D Model Retrieval
Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zan Gao 0002, Anan Liu
IEEE Trans. Multim.5
2025 CACP: Covariance-Aware Cross-Domain Prototypes for Domain Adaptive Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to reduce domain shifts / discrepancies between source and target domains, improving the source domain model's generalization ability to the target domain. Recently, prototypical methods, which primarily use single-source or single-target domain prototypes as category centers to aggregate features from both domains, have achieved competitive performance in this task. However, due to large domain shifts, single-source domain prototypes have finite generalization ability and not all source domain knowledge is conducive to model generalization. Single-target domain prototypes are noisy because they are prematurely initialized with all features filtered by pseudo labels, which causes error accumulation in the prototypes. To address these issues, we propose a covariance-aware cross-domain prototypes method (CACP) to achieve robust domain adaptation. We propose to use both domain prototypes to dynamically rectify pseudo labels in the target domain, effectively reducing the recognition difficulty of hard target domain samples and narrowing the gap between features of the same category in both domains. In addition, to further generalize the model to the target domain, we propose two modules based on covariance correlation, FSPC (Features Selection by Prototypes Covariances) and WSPC (Weighting Source by Prototypes Coefficients), to learn discriminative characteristics. FSPC selects highly correlated features to update target domain prototypes online, denoising and enhancing discriminativeness between categories. WSPC utilizes the correlation coefficients between target domain prototypes and source domain features to weight each point in the source domain, eliminating the information interference from the source domain. In particular, CACP achieves excellent performance on the GTA5$\to$Cityscapes and SYNTHIA$\to$Cityscapes tasks with minimal computational resources and time.
Yanbing Xue, Feifei Zhang 0001, Xianbin Wen, Zan Gao 0002, Shengyong Chen
IEEE Trans. Multim.5
2024 MS-YOLOv5s: An Improved YOLOv5s for the Detection of Imperceptible Defects on Steel Surfaces
Mian Zhou, Zan Gao 0002
ICIC (10)5
2023 Multi-Behavior Recommendation with Cascading Graph Convolution Networks
abstract
Multi-behavior recommendation, which exploits auxiliary behaviors (e.g., click and cart) to help predict users’ potential interactions on the target behavior (e.g., buy), is regarded as an effective way to alleviate the data sparsity or cold-start issues in recommendation. Multi-behaviors are often taken in certain orders in real-world applications (e.g., click>cart>buy). In a behavior chain, a latter behavior usually exhibits a stronger signal of user preference than the former one does. Most existing multi-behavior models fail to capture such dependencies in a behavior chain for embedding learning. In this work, we propose a novel multi-behavior recommendation model with cascading graph convolution networks (named MB-CGCN). In MB-CGCN, the embeddings learned from one behavior are used as the input features for the next behavior’s embedding learning after a feature transformation operation. In this way, our model explicitly utilizes the behavior dependencies in embedding learning. Experiments on two benchmark datasets demonstrate the effectiveness of our model on exploiting multi-behavior data. It outperforms the best baseline by 33.7% and 35.9% on average over the two datasets in terms of Recall@10 and NDCG@10, respectively.
Zhiyong Cheng 0001, Sai Han, Fan Liu 0008, Lei Zhu 0002, Zan Gao 0002, Yuxin Peng 0001
WWW5
2022 HMTN: Hierarchical Multi-scale Transformer Network for 3D Shape Recognition
abstract
As an important field of multimedia, 3D shape recognition has attracted much research attention in recent years. Various approaches have been proposed, within which the multiview-based methods show their promising performances. In general, an effective 3D shape recognition algorithm should take both the multiview local and global visual information into consideration, and explore the inherent properties of generated 3D descriptors to guarantee the performance of feature alignment in the common space. To tackle these issues, we propose a novel Hierarchical Multi-scale Transformer Network (HMTN) for the 3D shape recognition task. In HMTN, we propose a multi-level regional transformer (MLRT) module for shape descriptor generation. MLRT includes two branches that aim to extract the intra-view local characteristics by modeling region-wise dependencies and give the supervision of multiview global information under different granularities. Specifically, MLRT can comprehensively consider the relations of different regions and focus on the discriminative parts, which improves the effectiveness of the learned descriptors. Finally, we adopt the cross-granularity contrastive learning (CCL) mechanism for shape descriptor alignment in the common space. It can explore and utilize the cross-granularity semantic correlation to guide the descriptor extraction process while performing the instance alignment based on the category information. We evaluate the proposed network on several public benchmarks, and HMTN achieves competitive performance compared with the state-of-the-art (SOTA) methods.
Yue Zhao 0042, Weizhi Nie, Zan Gao 0002, Anan Liu
ACM Multimedia3
2022 Semantically guided projection for zero-shot 3D model classification and retrieval
Yuting Su 0001, Jiayu Li 0004, Wenhui Li 0001, Zan Gao 0002, Haipeng Chen 0002, Xuanya Li, Anan Liu
Multim. Syst.4
2022 Joint Local Correlation and Global Contextual Information for Unsupervised 3D Model Retrieval and Classification
abstract
Unsupervised 3D model analysis has attracted tremendous attentions with the increasing growth of 3D model data and the extensive human annotations. Many effective methods have been designed to address the 3D model analysis with labeled information, while rare methods devote to unsupervised deep learning due to the difficulty of mining reliable information. In this paper, we propose a novel unsupervised deep learning method named joint local correlation and global contextual information (LCGC) for 3D model retrieval and classification, which mines the reliable triplet set and uses triplet loss to optimize the deep neural network. Our method proposes two schemes: 1) Local self-correlation information learning, which adopts the intra and inter information to construct the view-level triplet set. 2) Global neighbor contextual information learning, which employs the neighbor contextual information to explore the reliable relations among 3D models and construct the model-level triplet set. The above schemes encourage that the selected triple set can been used to improve the discrimination of learned features. Extensive evaluations on two large-scale datasets, ModelNet40 and ShapeNet55, have demonstrated the effectiveness of our proposed method.
Wenhui Li 0001, Zhenlan Zhao, Anan Liu, Zan Gao 0002, Chenggang Yan 0001, Zhendong Mao 0001, Haipeng Chen 0002, Weizhi Nie
IEEE Trans. Circuits Syst. Video Technol.4
2022 Domain-Adversarial-Guided Siamese Network for Unsupervised Cross-Domain 3-D Object Retrieval
abstract
Recent advances in 3-D sensors and 3-D modeling have led to the availability of massive amounts of 3-D data. It is too onerous and time consuming to manually label a plentiful of 3-D objects in real applications. In this article, we address this issue by transferring the knowledge from the existing labeled data (e.g., the annotated 2-D images or 3-D objects) to the unlabeled 3-D objects. Specifically, we propose a domain-adversarial guided siamese network (DAGSN) for unsupervised cross-domain 3-D object retrieval (CD3DOR). It is mainly composed of three key modules: 1) siamese network-based visual feature learning; 2) mutual information (MI)-based feature enhancement; and 3) conditional domain classifier-based feature adaptation. First, we design a siamese network to encode both 3-D objects and 2-D images from two domains because of its balanced accuracy and efficiency. Besides, it can guarantee the same transformation applied to both domains, which is crucial for the positive domain shift. The core issue for the retrieval task is to improve the capability of feature abstraction, but the previous CD3DOR approaches merely focus on how to eliminate the domain shift. We solve this problem by maximizing the MI between the input 3-D object or 2-D image data and the high-level feature in the second module. To eliminate the domain shift, we design a conditional domain classifier, which can exploit multiplicative interactions between the features and predictive labels, to enforce the joint alignment in both feature level and category level. Consequently, the network can generate domain-invariant yet discriminative features for both domains, which is essential for CD3DOR. Extensive experiments on two protocols, including the cross-dataset 3-D object retrieval protocol (3-D to 3-D) on PSB/NTU, and the cross-modal 3-D object retrieval protocol (2-D to 3-D) on MI3DOR-2, demonstrate that the proposed DAGSN can significantly outperform state-of-the-art CD3DOR methods.
Anan Liu, Fu-Bin Guo, Heyu Zhou, Chenggang Yan 0001, Zan Gao 0002, Xuanya Li, Wenhui Li 0001
IEEE Trans. Cybern.5
2021 SVHAN: Sequential View Based Hierarchical Attention Network for 3D Shape Recognition
abstract
As an important field of multimedia, 3D shape recognition has attracted much research attention in recent years. A lot of deep learning models have been proposed for effective 3D shape representation. The view-based methods show the superiority due to the comprehensive exploration of the visual characteristics with the help of established 2D CNN architectures. Generally, the current approaches contain the following disadvantages: First, the most majority of methods lack the consideration for sequential information among the multiple views, which can provide descriptive characteristics for shape representation. Second, the incomprehensive exploration for the multi-view correlations directly affects the discrimination of shape descriptors. Finally, roughly aggregating multi-view features leads to the loss of descriptive information, which limits the shape representation effectiveness. To handle these issues, we propose a novel sequential view based hierarchical attention network (SVHAN) for 3D shape recognition. Specifically, we first divide the view sequence into several view blocks. Then, we introduce a novel hierarchical feature aggregation module (HFAM), which hierarchically exploits the view-level, block-level, and shape-level features, the intra- and inter- view-block correlations are also captured to improve the discrimination of learned features. Subsequently, a novel selective fusion module (SFM) is designed for feature aggregation, considering the correlations between different levels and preserving effective information. Finally, discriminative and informative shape descriptors are generated for the recognition task. We validate the effectiveness of our proposed method on two public databases. The experimental results show the superiority of SVHAN against the current state-of-the-art approaches.
Yue Zhao 0042, Weizhi Nie, Anan Liu, Zan Gao 0002, Yuting Su 0001
ACM Multimedia4
2021 Pairwise attention network for cross-domain image recognition
Zan Gao 0002, Guangping Xu, Xianbin Wen
Neurocomputing1
2021 A bus passenger re-identification dataset and a deep learning baseline using triplet embedding
Junliang Guo, Yanbing Xue, Zan Gao 0002, Guangping Xu, Hua Zhang 0003
Multim. Tools Appl.4
2020 Multi-graph Convolutional Network for Unsupervised 3D Shape Retrieval
abstract
3D shape retrieval has attracted much research attention due to its wide applications in the fields of computer vision and multimedia. Various approaches have been proposed in recent years for learning 3D shape descriptor from different modalities. The existing works contain the following disadvantages: 1) the vast majority methods rely on the large scale of training data with clear category information; 2) many approaches focus on the fusion of multi-modal information but ignore the guidance of correlations among different modalities for shape representation learning; 3) many methods pay attention to the structural feature learning of 3D shape but ignore the guidance of structural similarity between every two shapes. To solve these problems, we propose a novel multi-graph network (MGN) for unsupervised 3D shape retrieval, which utilizes the correlations among modalities and structural similarity between two models to guide the shape representation learning process without category information. More specifically, we propose two novel loss functions: auto-correlation loss and cross-correlation loss. The auto-correlation loss utilizes information from different modalities to increase the discrimination of shape descriptor. The cross-correlation loss utilizes the structural similarity between two models to strengthen the intra-class similarity and increase the inter-class distinction. Finally, an effective similarity measurement is designed for the shape retrieval task. To validate the effectiveness of our proposed method, we conduct experiments on the ModelNet dataset. Experimental results demonstrate the effectiveness of our proposed method, and significant improvements have been achieved compared with state-of-the-art methods.
Weizhi Nie, Yue Zhao 0042, Anan Liu, Zan Gao 0002, Yuting Su 0001
ACM Multimedia4
2020 Exploring Deep Learning for View-Based 3D Model Retrieval
abstract
In recent years, view-based 3D model retrieval has become one of the research focuses in the field of computer vision and machine learning. In fact, the 3D model retrieval algorithm consists of feature extraction and similarity measurement, and the robust features play a decisive role in the similarity measurement. Although deep learning has achieved comprehensive success in the field of computer vision, deep learning features are used for 3D model retrieval only in a small number of works. To the best of our knowledge, there is no benchmark to evaluate these deep learning features. To tackle this problem, in this work we systematically evaluate the performance of deep learning features in view-based 3D model retrieval on four popular datasets (ETH, NTU60, PSB, and MVRED) by different kinds of similarity measure methods. In detail, the performance of hand-crafted features and deep learning features are compared, and then the robustness of deep learning features is assessed. Finally, the difference between single-view deep learning features and multi-view deep learning features is also evaluated. By quantitatively analyzing the performances on different datasets, it is clear that these deep learning features can consistently outperform all of the hand-crafted features, and they are also more robust than the hand-crafted features when different degrees of noise are added into the image. The exploration of latent relationships among different views in multi-view deep learning network architectures shows that the performance of multi-view deep learning outperforms that of single-view deep learning features with low computational complexity.
Zan Gao 0002, Yinmin Li, Shaohua Wan 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Cognitive-inspired class-statistic matching with triple-constrain for camera free 3D object retrieval
Zan Gao 0002, Shaohua Wan 0001, Hua Zhang 0003, Yinglong Wang 0001
Future Gener. Comput. Syst.1
2018 3D object recognition based on pairwise Multi-view Convolutional Neural Networks
Zan Gao 0002, Yanbin Xue, Guangping Xu, Hua Zhang 0003, Yinglong Wang 0001
J. Vis. Commun. Image Represent.1
2018 MMA: a multi-view and multi-modality benchmark dataset for human action recognition
Zan Gao 0002, Tao-tao Han, Hua Zhang 0003, Yanbing Xue, Guangping Xu
Multim. Tools Appl.1
2017 Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantage
abstract
Segment-tree (ST) based cost aggregation algorithm for stereo matching successfully integrates the information of segmentation with non-local cost aggregation framework. The tree structure which is generated by the segmentation strategy directly determines the final results for this kind of algorithms. However, the original strategy performs unreasonable due to its coarse performance and ignores to meet the disparity consistency assumption. To improve these weaknesses we propose a novel segmentation algorithm for constructing a more faithful ST with enhanced segmentation advantage according to a robust initial over-segmentation. Then we implement non-local cost aggregation framework on this new ST structure and obtain improved disparity maps. Performance evaluations on all 31 Middlebury stereo pairs show that the proposed algorithm outperforms than other five state-of-the-art aggregated based algorithms and also keeps time efficiency.
Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002, Shengyong Chen
ICASSP6
2017 3D human action recognition model based on image set and regularized multi-task leaning
Zan Gao 0002, Guotai Zhang, Hua Zhang 0003, Yanbin Xue, Guangping Xu
Neurocomputing1
2016 Iterative color-depth MST cost aggregation for stereo matching
abstract
The minimum spanning tree (MST) based non-local cost aggregation algorithm performs well in accuracy and time efficiency. However, it can still be improved in two aspects. First, we propose a logarithmic transformation on matching cost function to improve the matching efficiency in texture less regions. The textureless neighbors can provide effective contributions in cost aggregation by the proposed monotone increasing function. Hence the algorithm can distinguish different pixels in textureless regions. Second, MST algorithm only utilizes color information in weight function while aggregating, which leads 3D cues missing. We introduce depth weight computed from the original MST algorithm into an edge weight function. With the proposed color-depth weight, we further iteratively rebuild the tree and obtain enhanced disparity map. Performance evaluations on 19 Middlebury stereo pairs and Microsoft stereo videos show that the proposed algorithm outperforms than other five state-of-the-art cost aggregation algorithms.
Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002
ICME6
2016 A Fast 3D Retrieval Algorithm via Class-Statistic and Pair-Constraint Model
abstract
With the development of 3D technologies and devices, 3D model retrieval becomes a hot research topic where multi-view matching algorithms have demonstrated satisfying performance. However, exciting works overlook the common factors among objects in a single class, and they are time consuming in retrieval processing. In this paper, a class-statistics and pair-constraint model (CSPC) method is originally proposed for 3D model retrieval, which is composed of supervised class-based statistics model and pair-constraint object retrieval model. In our CSPC model, we firstly convert view-based distance measure into object-based distance measure without falling in performance, which will advance 3D model retrieval speed. Secondly, the generality of the distribution of each feature dimension in each class is computed to judge category information, and then we further adopt this distribution information to build class models. Finally, an object-based pairwise constraint is introduced on the base of the class-statistic measure, which can remove a lot of false alarm samples in retrieval. Experimental results on ETH, NTU-60, MVRED and PSB 3D datasets show that our method is fast, and its performance is also comparable with the-state-of-the-art algorithms.
Zan Gao 0002, Hua Zhang 0003, Yanbing Xue, Guangping Xu
ACM Multimedia1
2016 Evaluation of local spatial-temporal features for cross-view action recognition
Zan Gao 0002, Weizhi Nie, Anan Liu, Hua Zhang 0003
Neurocomputing1
2016 Human action recognition on depth dataset
Zan Gao 0002, Hua Zhang 0003, Anan Liu, Guangping Xu, Yanbing Xue
Neural Comput. Appl.1
2015 Clique-graph matching by preserving global & local structure
abstract
This paper originally proposes the clique-graph and further presents a clique-graph matching method by preserving global and local structures. Especially, we formulate the objective function of clique-graph matching with respective to two latent variables, the clique information in the original graph and the pairwise clique correspondence constrained by the one-to-one matching. Since the objective function is not jointly convex to both latent variables, we decompose it into two consecutive steps for optimization: 1) clique-to-clique similarity measure by preserving local unary and pairwise correspondences; 2) graph-to-graph similarity measure by preserving global clique-to-clique correspondence. Extensive experiments on the synthetic data and real images show that the proposed method can outperform representative methods especially when both noise and outliers exist.
Weizhi Nie, Anan Liu, Zan Gao 0002, Yuting Su 0001
CVPR3
2015 Single Face Image Super-Resolution via Multi-dictionary Bayesian Non-parametric Learning
Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002
ICONIP (1)6
2015 Multi-perspective and multi-modality joint representation and recognition model for 3D action recognition
Zan Gao 0002, Hua Zhang 0003, Guangping Xu, Yanbin Xue
Neurocomputing1
2015 Multi-view discriminative and structured dictionary learning with group sparsity for human action recognition
Zan Gao 0002, Hua Zhang 0003, Guangping Xu, Yanbin Xue, Alex Hauptmann 0001
Signal Process.1
2014 Enhanced and hierarchical structure algorithm for data imbalance problem in semantic extraction under massive video dataset
Zan Gao 0002, Ming-yu Chen 0001, Alex Hauptmann 0001, Hua Zhang 0003, Anni Cai
Multim. Tools Appl.1
2013 An Effective Tracking System for Multiple Object Tracking in Occlusion Scenes
Weizhi Nie, Anan Liu, Yuting Su 0001, Zan Gao 0002
MMM (1)4
2013 Online Boosting Tracking with Fragmented Model
Dingcheng Shen, Hua Zhang 0003, Yanbing Xue, Guangping Xu, Zan Gao 0002
MMM (2)5
2012 Human action recognition based on sparse representation induced by L1/L2 regulations
Zan Gao 0002, Anan Liu, Hua Zhang 0003, Guangping Xu, Yanbing Xue
ICPR1