Heyu Zhou

dblp:223/1857 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
11since 2021 · last 2024
0000-0001-9451-5600ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Multi-Modal Meta-Transfer Fusion Network for Few-Shot 3D Model Classification
Heyu Zhou, Anan Liu, Chenyu Zhang 0003, Qianyi Zhang, Mohan Kankanhalli
Int. J. Comput. Vis.1
2024 View sequence prediction GAN: unsupervised representation learning for 3D shapes by decomposing view content and viewpoint variance
Heyu Zhou, Jiayu Li 0004, Xianzhu Liu, Yingda Lyu, Haipeng Chen 0002, Anan Liu
Multim. Syst.1
2024 Universal unsupervised cross-domain 3D shape retrieval
Heyu Zhou, Qipei Liu, Jiayu Li 0004, Xuanya Li, Anan Liu
Multim. Syst.1
2024 Balanced Class-Incremental 3D Object Classification and Retrieval
abstract
Most existing 3D object classification and retrieval algorithms rely on one-off supervised learning on closed 3D object sets and tend to provide rigid convolutional neural networks with little scalability. Such limitations substantially restrict their potential to learn newly emerged 3D object classes continually in the real world. Aiming to go beyond these limitations, we innovatively propose two new and challenging tasks: class-incremental 3D object classification (CI-3DOC) and class-incremental 3D object retrieval (CI-3DOR), the key to which is class-incremental 3D representation learning. It expects the network to update continually to learn new 3D class representations without forgetting the previously learned ones. To this end, we design a novel balanced distillation network(BDNet)that uses a dual supervision mechanism to balance between consolidating old knowledge (stability) and adapting to new 3D object classes (plasticity) carefully. On the one hand, we employ stability-based supervision to retain the stable and discriminative information of old classes that greatly benefit both classification and retrieval tasks. On the other hand, we use plasticity-based supervision to improve the network's generalization for learning new class 3D representations by transferring knowledge from a temporary teacher network to the current model. By properly handling the relationship between the two modules, we achieve a surprising performance improvement. Furthermore, considering there is no available dataset for evaluation, we build two 3D datasets, INOR-1 and INOR-2, to evaluate these two new tasks. Extensive experimental results demonstrate that our method can significantly outperform other state-of-the-art class-incremental learning methods. Even if we store 500-1000 fewer 3D objects than SOTA methods,BDNetstill achieves comparable performance.
Anan Liu, Haochun Lu, Heyu Zhou, Tianbao Li 0001, Mohan Kankanhalli
IEEE Trans. Knowl. Data Eng.3
2023 Vulnerability of Feature Extractors in 2D Image-Based 3D Object Retrieval
abstract
Recent advances in 3D modeling software and 3D capture devices contribute to the availability of large-scale 3D objects. Together with the prevalence of deep neural networks (DNNs), DNN-based 3D object retrieval systems are widely applied, especially by inputting 2D images to retrieve 3D objects. Although DNNs have shown vulnerable to adversarial attacks in classification, the vulnerability of DNN-based 3D object retrieval system remains under-explored. In this paper, we formulate the problem of attacking against DNN-based feature extractors in the 2D image-based 3D object retrieval system. Specifically, we consider the attack happens under a reasonable scenario that the candidate 3D object database is unknown to the adversary, which challenges adversarial example generation. To tackle this difficulty, we set up a reasonable hypothesis on the information which the adversary can be accessible, and then propose two effective perturbation generation methods: one is to corrupt domain-level alignment (CDA) and the other one is to corrupt class-level alignment (CCA). In converse, we propose a novel progressive adversarial training (PAT) method to improve the feature extractor robustness, which can effectively and stably mitigate both CDA and CCA attacks. Experimental results demonstrate that a typical feature extractor can be effectively compromised by attacks. Moreover, the transferability of the adversarial query illustrates the possibility of realistic black-box attacks. The successful defense against both CDA and CCA attacks by PAT can validate the superiority of the proposed defense method.
Anan Liu, Heyu Zhou, Xuanya Li, Lanjun Wang
IEEE Trans. Multim.2
2022 Collaborative Distribution Alignment for 2D image-based 3D shape retrieval
Nian Hu, Heyu Zhou, Anan Liu, Xiangdong Huang 0002, Shenyuan Zhang, Guoqing Jin, Junbo Guo, Xuanya Li
J. Vis. Commun. Image Represent.2
2022 A Feature Transformation Framework With Selective Pseudo-Labeling for 2D Image-Based 3D Shape Retrieval
abstract
2D image-based 3D shape retrieval (2D-to-3D) aims at searching the corresponding 3D shapes (unlabeled) when given a 2D image (labeled), which is a fundamental task in computer vision and has gained a surge of attention in recent years. However, extensive prior works are limited by two settings, 1) reducing domain discrepancy while ignoring the 3D shape style, 2) 3D shapes are simply and brutally pseudo-annotated by the 2D image-supervised classifier, neglecting the structure information underlying the 3D shape domain. To remedy these issues, we propose a feature transformation framework with selective pseudo-labeling (FTSPL) for 2D-to-3D task. Specifically, we first employ CNNs to produce both 2D image and 3D shape (described as multiple views) features, then we force the inter-domain centroid alignment class-wisely to reduce the overall domain discrepancy. In addition to this, we further exploit the intra-category attribute variation (covariance) of 3D shape features to transform the 2D image features. By doing so, we can equip 2D features with 3D shape style. Since the centroid and covariance estimation of 3D shape features require accurate label predictions, we put forward a selective pseudo-labeling module, which can assign reliable pseudo-labels for 3D shapes via nearest category centroid and cluster analysis, respectively, while preserving the structure information of 3D shapes. Comprehensive experiments validate that our model surpasses the state-of-the-arts on standard 2D-to-3D benchmarks (MI3DOR and MI3DOR-2).
Nian Hu, Heyu Zhou, Xiangdong Huang 0002, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2022 Learning Transferable and Discriminative Representations for 2D Image-Based 3D Model Retrieval
abstract
Existing research on the 2D image-based 3D model retrieval task focuses on learning transferable representations directly to narrow the domain discrepancy. However, it is not easy to achieve in practice due to the significant variations across two domains. In addition, some methods design a domain discriminator to distinguish the feature arising from source or target domains for transferable feature representations learning, which will lead to an unexpected deterioration of the feature discriminability. To settle these problems, we propose jointly learning transferable and discriminative representations for 2D image-based 3D model retrieval. Specifically, we extract features from the 2D images and 3D models (described as multiple views) by CNN. Considering the difficulty of directly narrowing the discrepancy of two domains, we are prone to connect 2D image and 3D model domains to an intermediate domain, where the domain gap aims to be eliminated. However, the feature transferability does not denote well discriminability. Based on the batch spectral penalization (BSP) theory, the feature transferability is dominated by feature vectors with higher singular values, while the feature discriminability depends on more eigenvectors with lower singular values to convey rich discriminative structures. Therefore, we penalize the largest singular values so that the feature vectors with lower singular values are appropriately enhanced, thereby strengthening feature discriminability. A series of experiments on two challenging datasets, MI3DOR and MI3DOR-2, indicate that our method can significantly improve performance.
Yaqian Zhou 0002, Yu Liu 0004, Heyu Zhou, Zhiyong Cheng 0001, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2022 Domain-Adversarial-Guided Siamese Network for Unsupervised Cross-Domain 3-D Object Retrieval
abstract
Recent advances in 3-D sensors and 3-D modeling have led to the availability of massive amounts of 3-D data. It is too onerous and time consuming to manually label a plentiful of 3-D objects in real applications. In this article, we address this issue by transferring the knowledge from the existing labeled data (e.g., the annotated 2-D images or 3-D objects) to the unlabeled 3-D objects. Specifically, we propose a domain-adversarial guided siamese network (DAGSN) for unsupervised cross-domain 3-D object retrieval (CD3DOR). It is mainly composed of three key modules: 1) siamese network-based visual feature learning; 2) mutual information (MI)-based feature enhancement; and 3) conditional domain classifier-based feature adaptation. First, we design a siamese network to encode both 3-D objects and 2-D images from two domains because of its balanced accuracy and efficiency. Besides, it can guarantee the same transformation applied to both domains, which is crucial for the positive domain shift. The core issue for the retrieval task is to improve the capability of feature abstraction, but the previous CD3DOR approaches merely focus on how to eliminate the domain shift. We solve this problem by maximizing the MI between the input 3-D object or 2-D image data and the high-level feature in the second module. To eliminate the domain shift, we design a conditional domain classifier, which can exploit multiplicative interactions between the features and predictive labels, to enforce the joint alignment in both feature level and category level. Consequently, the network can generate domain-invariant yet discriminative features for both domains, which is essential for CD3DOR. Extensive experiments on two protocols, including the cross-dataset 3-D object retrieval protocol (3-D to 3-D) on PSB/NTU, and the cross-modal 3-D object retrieval protocol (2-D to 3-D) on MI3DOR-2, demonstrate that the proposed DAGSN can significantly outperform state-of-the-art CD3DOR methods.
Anan Liu, Fu-Bin Guo, Heyu Zhou, Chenggang Yan 0001, Zan Gao 0002, Xuanya Li, Wenhui Li 0001
IEEE Trans. Cybern.3
2021 Hierarchical multi-view context modelling for 3D object classification and retrieval
Anan Liu, Heyu Zhou, Weizhi Nie, Zhenguang Liu, Wu Liu 0005, Hongtao Xie 0001, Zhendong Mao 0001, Xuanya Li, Dan Song 0006
Inf. Sci.2
2021 Wasserstein distance feature alignment learning for 2D image-based 3D model retrieval
Yaqian Zhou 0002, Yu Liu 0004, Heyu Zhou, Wenhui Li 0001
J. Vis. Commun. Image Represent.3
2020 Hierarchical Instance Feature Alignment for 2D Image-Based 3D Shape Retrieval
abstract
2D image-based 3D shape retrieval has become a hot research topic since its wide industrial applications and academic significance. However, existing view-based 3D shape retrieval methods are restricted by two settings, 1) learn the common-class features while neglecting the instance visual characteristics, 2) narrow the global domain variations while ignoring the local semantic variations in each category. To overcome these problems, we propose a novel hierarchical instance feature alignment (HIFA) method for this task. HIFA consists of two modules, cross-modal instance feature learning and hierarchical instance feature alignment. Specifically, we first use CNN to extract both 2D image and multi-view features. Then, we maximize the mutual information between the input data and the high-level feature to preserve as much as visual characteristics of an individual instance. To mix up the features in two domains, we enforce feature alignment considering both global domain and local semantic levels. By narrowing the global domain variations we impose the identical large norm restriction on both 2D and 3D feature-norm expectations to facilitate more transferable possibility. By narrowing the local variations we propose to minimize the distance between two centroids of the same class from different domains to obtain semantic consistency. Extensive experiments on two popular and novel datasets, MI3DOR and MI3DOR-2, validate the superiority of HIFA for 2D image-based 3D shape retrieval task.
Heyu Zhou, Weizhi Nie, Wenhui Li 0001, Dan Song 0006, Anan Liu
IJCAI1
2020 Semantic Consistency Guided Instance Feature Alignment for 2D Image-Based 3D Shape Retrieval
abstract
2D image-based 3D shape retrieval (2D-to-3D) investigates the problem of matching the relevant 3D shapes from gallery dataset when given a query image. Recently, adversarial training and environmental style transfer learning have been successful applied to this task and achieved state-of-the-art performance. However, there still exist two problems. First, previous works only concentrate on the connection between the label and representation, where the unique visual characteristics of each instance are paid less attention. Second, the confused features or the transformed images can only cheat the discriminator but can not guarantee the semantic consistency. In another words, features of 2D desk may be mapped nearby the features of 3D chair. In this paper, we propose a novel semantic consistency guided instance feature alignment network (SC-IFA) to address these limitations. SC-IFA mainly consists of two parts, instance visual feature extraction and cross-domain instance feature adaptation. For the first module, unlike previous methods, which merely employ 2D CNN to extract the feature, we additionally maximize the mutual information between the input and feature to enhance the capability of feature representation for each instance. For the second module, we first introduce the margin disparity discrepancy model to mix up the cross-domain features in an adversarial training way. Then, we design two feature translators to transform the feature from one domain to another domain, and impose the translation loss and correlation loss on the transformed features to preserve the semantic consistency. Extensive experimental results on two benchmarks, MI3DOR and MI3DOR-2, verify SC-IFA is superior to the state-of-the-art methods.
Heyu Zhou, Weizhi Nie, Dan Song 0006, Nian Hu, Xuanya Li, Anan Liu
ACM Multimedia1
2020 3D model retrieval based on multi-view attentional convolutional neural network
Anan Liu, Heyu Zhou, Meng-Jie Li, Weizhi Nie
Multim. Tools Appl.2
2020 Multi-View Saliency Guided Deep Neural Network for 3-D Object Retrieval and Classification
abstract
In this paper, we propose the multi-view saliency guided deep neural network (MVSG-DNN) for 3D object retrieval and classification. This method mainly consists of three key modules. First, the module of model projection rendering is employed to capture the multiple views of one 3D object. Second, the module of visual context learning applies the basic Convolutional Neural Networks for visual feature extraction of individual views and then employs the saliency LSTM to adaptively select the representative views based on multi-view context. Finally, with these information, the module of multi-view representation learning can generate the compile 3D object descriptors with the designed classification LSTM for 3D object retrieval and classification. The proposed MVSG-DNN has two main contributions: 1) It can jointly realize the selection of representative views and the similarity measure by fully exploiting multi-view context; 2) It can discover the discriminative structure of multi-view sequence without constraints of specific camera settings. Consequently, it can support flexible 3D object retrieval and classification for real applications by avoiding the required camera settings. Extensive comparison experiments on ModelNet10, ModelNet40, and ShapeNetCore55 demonstrate the superiority of MVSG-DNN against the state-of-art methods.
Heyu Zhou, Anan Liu, Weizhi Nie, Jie Nie
IEEE Trans. Multim.1
2019 Dual-level Embedding Alignment Network for 2D Image-Based 3D Object Retrieval
abstract
Recent advances in 3D modeling software and 3D capture devices contribute to the availability of large-scale 3D objects. However, manually labelled large-scale 3D object dataset is still too expensive to build in practice. An intuitive idea is to transfer the knowledge from label-rich 2D images (source domain) to unlabelled 3D objects (target domain) to facilitate 3D big data management. In this paper, we propose an unsupervised dual-level embedding alignment (DLEA) network for a new task, 2D image-based 3D object retrieval. It mainly consists of two modules, visual feature learning and cross-domain feature adaptation, for jointly optimizing. The first module transforms individual 3D object into a set of multi-view images and utilizes 2D CNNs to extract visual features of both multi-view image sets and the source 2D images. For multi-view fusion by reducing the distribution divergence between both domains, we propose a cross-domain view-wise attention mechanism to adaptively compute the weights of individual views and aggregate them into a compact descriptor to narrow the gap between source and target domains. With the visual representation of both domains, the module of cross-domain feature adaptation aims to enforce the domain-level and class-level embedding alignment of cross-domain feature spaces. For domain-level embedding alignment, we train a discriminator to align the global distribution statistics of both spaces. For class-level embedding alignment, we map the features in the same class but from different domains nearby through aligning the centroid of each class from both domains. To our knowledge, this is the first unsupervised work to jointly realize cross-domain feature learning and distribution alignment in an end-to-end manner for this new task. Moreover, we constructed two new datasets, MI3DOR and MI3DOR-2, to advocate the research on this topic. Extensive comparison experiments can demonstrate the superiority of DLEA against the state-of-art methods.
Heyu Zhou, Anan Liu, Weizhi Nie
ACM Multimedia1