ChengKai Xia

dblp:416/5785 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › point cloud processing
point cloud completion
0.912025
Points, Images and Texts: Boosting Point Cloud Completion with Multi-Modal Features · ICRA 2025
Computer vision › 3D vision › 3d shape modeling
shape refinement
0.912025
Points, Images and Texts: Boosting Point Cloud Completion with Multi-Modal Features · ICRA 2025

Methods — techniques the papers use, named apart from their topics

visual-textual embedding · 0.9visual question answering · 0.9multi-scale edge convolution · 0.9cross-attention · 0.9
YearPublicationVenuePosition
2025 Points, Images and Texts: Boosting Point Cloud Completion with Multi-Modal Features
abstract
Point cloud completion is crucial for reconstructing accurate shapes in many 3D visual applications. Recent approaches incorporate images into the completion pipeline, introducing geometric clues and global constraints. However, their fusion processes often fail to reconstruct detailed parts and maintain global consistency simultaneously. Except for images, text is another important clue for recognizing the target's characteristics. Thus, in this work, we propose to combine multiple modalities including points, images and texts for point cloud completion. Specifically, inspired by recently pre-trained large language models, we generate the description texts for images by Visual Question Answering (VQA) models and introduce Visual-Textual Embedding (VTE) models to extract joint features of image-text pairs. Furthermore, we describe the edge geometric patterns by multi-scale edge convolution to guide the refinement of shapes in local areas. Then we adopt cross attention mechanism to effectively fuse multi-modal features and refine the coarse shape. Extensive experiments on commonly used benchmarks demonstrate our method's superior performance over previous uni-modal and cross-modal methods.
ChengKai Xia, Fan Lu 0001, Bin Li 0087, Alois C. Knoll, Guang Chen 0001
ICRA1