Qi Liang 0004

dblp:147/6275-4 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0001-5598-6012ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SCDL: Sketch Causal Disentangled Learning for Sketch-Based 3D Shape Retrieval
abstract
Sketch-based 3D shape retrieval (SBSR) has been a challenging task for decades, crucially depending on aligning shared semantic attributes between sketches and 3D shapes. Previous efforts mainly aimed at creating a common embedding space to bridge domain gaps. However, sketches’ subjective and abstract nature, known as confounders, potentially reduces learning performance of matching with 3D shapes. To address this issue, in this paper, we propose a sketch causal disentangled learning for SBSR, named SCDL, which introduce causal intervention to explicitly disentangle sketches into the inherent shared semantic part, and other unrelated confounders to classification (styles, abstraction levels, etc.) for the first time. Specifically, we construct a structural causal model (SCM) in the sketch branch under the dual variational autoencoder (VAE) architectures to alleviate confounders negative impact through learning the semantic attributes in the latent variable space. Next, we adopt a learning strategy on the separated semantic latent variables to construct a shared semantic embedding space further to make cross-modal features of the same class more similar, alleviating the cross-modality discrepancies effectively and establishing new state-of-the-art on three benchmarks. Comprehensive experiment results, ablation studies, and visualization validate the effectiveness of our approach.
Shaojin Bai, Yalu Li, Rihao Chang, Qi Liang 0004, Weizhi Nie
IEEE Trans. Circuits Syst. Video Technol.4
2023 Principal views selection based on growing graph convolution network for multi-view 3D model recognition
Qi Liang 0004, Qiang Li 0048, Weizhi Nie, Yuting Su 0001
Appl. Intell.1
2023 Unsupervised Cross-Media Graph Convolutional Network for 2D Image-Based 3D Model Retrieval
abstract
With the rapid development of 3D construction technology, 3D models have been implemented in many applications. In particular, the fields of virtual and augmented reality have created a considerable demand for rapid access to large sets of 3D models in recent years. An effective method for addressing the demand is to search 3D models based on 2D images because 2D images can be easily captured by smartphones or other lightweight vision sensors. In this paper, we propose a novel unsupervised cross-media graph convolutional network (UCM-GCN) for 3D model retrieval based on 2D images. Here, we render views from 3D models to construct a graph model based on 3D model structural information. Then, we utilize the 2D image's visual information to bridge the gap between cross-modality data. Then, the proposed UCM-GCN is utilized to update the feature vector of the 2D image and the 3D model. Here, we introduce correlation loss to mitigate the distribution discrepancy across different modalities, which can fully consider the structural and visual similarities between the 2D image and 3D model to embed the final different modalities into the same feature space. To demonstrate the performance of our approach, we conducted a series of experiments on the MI3DOR dataset, which is utilized in SHREC19. We also compared it with other similar methods on the 3D-FUTURE dataset. The experimental results demonstrate the superiority of our proposed method over state-of-the-art methods.
Qi Liang 0004, Qiang Li 0048, Weizhi Nie, Anan Liu
IEEE Trans. Multim.1
2022 LP-GAN: Learning perturbations based on generative adversarial networks for point cloud adversarial attacks
Qi Liang 0004, Qiang Li 0048
Image Vis. Comput.1
2022 JFLN: Joint Feature Learning Network for 2D sketch based 3D shape retrieval
Yue Zhao 0042, Qi Liang 0004, Ruixin Ma, Weizhi Nie, Yuting Su 0001
J. Vis. Commun. Image Represent.2
2022 PAGN: perturbation adaption generation network for point cloud adversarial defense
Qi Liang 0004, Qiang Li 0048, Weizhi Nie, Anan Liu
Multim. Syst.1
2022 LD-GAN: Learning perturbations for adversarial defense based on GAN structure
Qi Liang 0004, Qiang Li 0048, Weizhi Nie
Signal Process. Image Commun.1
2021 3D shape recognition based on multi-modal information fusion
Qi Liang 0004, Mengmeng Xiao, Dan Song 0006
Multim. Tools Appl.1
2021 Exposing DeepFake Videos Using Attention Based Convolutional LSTM Network
Yishan Su, Huawei Xia, Qi Liang 0004, Weizhi Nie
Neural Process. Lett.3
2021 MHFP: Multi-view based hierarchical fusion pooling method for 3D shape recognition
Qi Liang 0004, Qiang Li 0048, Lihu Zhang, Haixiao Mi, Weizhi Nie, Xuanya Li
Pattern Recognit. Lett.1
2021 MMFN: Multimodal Information Fusion Networks for 3D Model Classification and Retrieval
abstract
In recent years, research into 3D shape recognition in the field of multimedia and computer vision has attracted wide attention. With the rapid development of deep learning, various deep models have achieved state-of-the-art performance based on different representations. There are many modalities for representing a 3D model, such as point cloud, multiview, and panorama view. Deep learning models based on these different modalities have different concerns, and all of them have achieved high performance for 3D shape recognition. However, all of these methods ignore the multimodality information in conditions where the same 3D model is represented by different modalities. Thus, we can obtain a better descriptor by guiding the training to consider these multiple representations. In this article, we propose MMFN, a novel multimodal fusion network for 3D shape recognition that employs correlations between the different modalities to generate a fused descriptor, which is more robust. In particular, we design two novel loss functions to help the model learn the correlation information during training. The first is correlation loss, which focuses on the correlations among different descriptors generated from different structures. This approach reduces the training time and improves the robustness of the fused descriptor of the 3D model. The second is instance loss, which preserves the independence of each modality and utilizes feature differentiation to guide model learning during the training process. More specifically, we use the weighted fusion method, which applies statistical methods to obtain robust descriptors that maximize the advantages of the information from the different modalities. We evaluated the proposed method on the ModelNet40 and ShapeNetCore55 datasets for 3D shape classification and retrieval tasks. The experimental results and comparisons with state-of-the-art methods demonstrate the superiority of our approach.
Weizhi Nie, Qi Liang 0004, Yuting Su 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2019 MMJN: Multi-Modal Joint Networks for 3D Shape Recognition
abstract
3D shape recognition has attracted wide research attention in the field of multimedia and computer vision. With the recent advance of deep learning, various deep models with different representations have achieved the state-of-the-art performances. Among them, many modalities are proposed to represent 3D model, such as point cloud, multi-view, and PANORAMA-view. Based on these representations, many corresponding deep models have shown significant performances on 3D shape recognition. However, few work to considers utilizing the fusion information of multi-modal for 3D shape recognition. Since these different modalities represent the same 3D model, they should guide each other to get a better feature representation. In this paper, we propose a novel multi-modal joint network (MMJN) for 3D shape recognition, which can consider the correlation between two different modalities to extract the robust feature vector. More specifically, we propose a novel correlation loss which can utilize the correlation between different features extracted by different modality networks to increase the robustness of the feature representation. Finally, we utilize the late fusion method to fuse the multi-modal information for 3D model representation and recognition. Here, we define the weight of different modalities features based on the statistic method and utilize the advantages of different modalities to generate more robust feature. We evaluated the proposed method on the ModelNet40 dataset for 3D shape classification and retrieval tasks. Experimental results and comparisons with the state-of-the-art methods demonstrate the superiority of our approach.
Weizhi Nie, Qi Liang 0004, Anan Liu, Zhendong Mao 0001
ACM Multimedia2
2019 Panorama based on multi-channel-attention CNN for 3D model recognition
Weizhi Nie, Qi Liang 0004, Roubing He
Multim. Syst.3