Jiayu Li 0004

dblp:147/0314-4 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-3458-2884ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Prompt-Tuning-Guided Dual-Distribution Alignment for Unsupervised 2D-3D Cross-Modal Retrieval
abstract
2D-3D cross-modal retrieval (2D-3DCMR) aims at retrieving the most matching 3D models by leveraging 2D query images. However, the inherent modal discrepancy makes the 2D-3DCMR task still largely challenging. Besides, the scarcity of 3D labels in real-world applications severely hinders learning discriminative representations. To address these limitations, we present a prompt tuning guided dual distribution alignment (PTG-DDA) framework based on the CLIP model for the 2D-3DCMR task. Specifically, we design a learnable multi-view adaptive representation learning (MARL) module that adaptively integrates 3D features to merge complementary information and filter out redundant information across views, thereby improving the representation capability of 3D models. To mitigate the feature distribution shift between the 2D and 3D data, we design an attention-guided heterogeneous feature alignment (AHFA) module to guide the 2D and 3D inputs attend to feature banks by adopting the attention mechanism, thereby achieving heterogeneous feature alignment. Furthermore, to learn discriminative 3D features, we employ a multi-modal semantic prompt synergy (MSPS) module, which integrates class-related representations into learnable prompts to progressively learn the cross-modal synergy via a prompt synergy adapter, thereby achieving semantic feature alignment. Comprehensive experimental results on popular 2D-3DCMR benchmarks, i.e., MI3DOR and MI3DOR-2, demonstrate the superiority and effectiveness of PTG-DDA.
Yaqian Zhou 0002, Ruiqiang Guo, Dan Song 0006, Jiayu Li 0004, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.4
2026 Separating Domain-Private Classes for Universal Unsupervised Cross-Domain 3D Model Retrieval
Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zan Gao 0002, Anan Liu
IEEE Trans. Multim.1
2025 Progressive Contrastive Label Optimization for Source-Free Universal 3D Model Retrieval
abstract
Unsupervised Cross-Domain 3D Model Retrieval (UCD3DMR) has emerged as an effective tool for managing 3D model data recently. However, existing UCD3DMR algorithms typically demand accessibility to source data and cross-domain label consistency, limiting their deployment in real-world industrial scenarios. Therefore, we relax the two demanding constraints and explore to address a newly challenging task, source-free universal 3D model retrieval (SFU3DMR). However, the inaccessibility to source data results in significant label noise in target pseudo-labels, while cross-domain label inconsistency introduces interference from target-private models, presenting tremendous challenges to model transfer. To address these challenges, we propose a novel SFU3DMR algorithm, Progressive Contrastive Label Optimization (PCLO). Specifically, we introduce the Neighbor-based Soft Label Optimization (NSLO) strategy, which refines target pseudo-labels based on the pseudo-label confidence of their nearest neighbors. Additionally, we design the Adaptive Hybrid Label Optimization (AHLO) strategy, which conducts positive label optimization to maximize label semantics for target-common models and executes negative label optimization to minimize label noise for target-private models. Experimental results confirm that the combined NSLO and AHLO strategies effectively refine the target pseudo-labels, and our PCLO achieves state-of-the-art performance for SFU3DMR on two well-established cross-domain benchmarks (MI3DOR and NTU/PSB).
Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, You Yang 0002, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.1
2024 View sequence prediction GAN: unsupervised representation learning for 3D shapes by decomposing view content and viewpoint variance
Heyu Zhou, Jiayu Li 0004, Xianzhu Liu, Yingda Lyu, Haipeng Chen 0002, Anan Liu
Multim. Syst.2
2024 Universal unsupervised cross-domain 3D shape retrieval
Heyu Zhou, Qipei Liu, Jiayu Li 0004, Xuanya Li, Anan Liu
Multim. Syst.4
2023 Cross-domain Prototype Contrastive loss for Few-shot 2D Image-Based 3D Model Retrieval
abstract
2D image-based 3D model retrieval (IBMR) usually relies on abundant explicit supervision on 2D images, together with unlabeled 3D models to learn domain-aligned yet class-discriminative features for the retrieval task. However, collecting large-scale 2D labels is cost-effective and time-consuming. Therefore, we explore a challenging IBMR task, where only few-shot labeled 2D images are available while the rest of the 2D and 3D samples remain unlabeled. Limited annotation of 2D images further increases the difficulty of domain-aligned yet discriminative feature learning. Therefore, we propose cross-domain prototype contrastive loss (CPCL) for the few-shot IBMR task. Specifically, we capture semantic information to learn class-discriminative features in each domain by minimizing intra-domain prototype contrastive loss. Besides, we perform inter-domain transferable contrastive learning to align the features between instances and prototypes of the same class across domains. Comprehensive experiments on popular benchmarks, MI3DOR and MI3DOR-2, validate the superiority of CPCL.
Yaqian Zhou 0002, Yu Liu 0004, Dan Song 0006, Jiayu Li 0004, Xuanya Li, Anan Liu
ICME4
2022 Semantically guided projection for zero-shot 3D model classification and retrieval
Yuting Su 0001, Jiayu Li 0004, Wenhui Li 0001, Zan Gao 0002, Haipeng Chen 0002, Xuanya Li, Anan Liu
Multim. Syst.2