EDBT 2026 Demo / reviewers in the wild / expert
Qiao Yu 0002
dblp:162/6793-2
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-6392-9461ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Point Cloud Representation Learning for Freeing Transformer to Perceive LocalabstractTransformers are widely utilized in the point cloud domain. However, existing methods tend to overburden Transformer with the dual task of local geometric perception and global feature extraction, limiting its ability to capture highlevel semantic knowledge. To address this issue, we present Representation Decoder (R-Decoder), a novel representation extraction module compatible with various point cloud Transformer methods, enabling the Transformer to focus on its excellent local perception. The R-Decoder iteratively extracts multiple global features from tokens generated by Transformer, refining them to construct an overall representation of point cloud. To ensure full adaptation of the R-Decoder to the knowledge of pre-trained Transformers, we design a cross-modal representation alignment task that leverages multimodal knowledge to specifically pre-train the R-Decoder. As a post-processing module, the R-Decoder seamlessly integrates with Transformers, while decoupling local perception and global representation. This design allows the Transformer to focus on the semantic encoding role for point tokens. Extensive experiments show that our RDecoder significantly boosts the capabilities of 3D representation learning in various point cloud Transformer methods. Notably, it achieves impressive classification accuracies of 95.1% on the ScanObjectNN dataset and 95.3% on the ModelNet40 dataset. Moreover, our method obtains new SOTA on all benchmarks of few-shot and zero-shot classification, while enhancing the multimodal task capabilities of pre-trained Transformers. Code and weights are available athttps://github.com/TangYuan96/RDecoder. Yunlong Yu 0002, Xianzhi Li 0001, Rui Wang 0077, Jinfeng Xu 0002, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
IEEE Trans. Multim. | 6 |
| 2025 | More Text, Less Point: Towards 3D Data-Efficient Point-Language UnderstandingabstractEnabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs has yet to be replicated in 3D understanding. In this paper, we rethink this issue and propose a new task: 3D Data-Efficient Point-Language Understanding. The goal is to enable LLMs to achieve robust 3D object understanding with minimal 3D point cloud and text data pairs. To address this task, we introduce GreenPLM, which leverages more text data to compensate for the lack of 3D data. First, inspired by using CLIP to align images and text, we utilize a pre-trained point cloud-text encoder to map the 3D point cloud space to the text space. This mapping leaves us to seamlessly connect the text space with LLMs. Once the point-text-LLM connection is established, we further enhance text-LLM alignment by expanding the intermediate text space, thereby reducing the reliance on 3D point cloud data. Specifically, we generate 6M free-text descriptions of 3D objects, and design a three-stage training strategy to help LLMs better explore the intrinsic connections between different modalities. To achieve efficient modality alignment, we design a zero-parameter cross-attention module for token pooling. Extensive experimental results show that GreenPLM requires only 12% of the 3D training data used by existing state-of-the-art models to achieve superior 3D understanding. Remarkably, GreenPLM also achieves competitive performance using text-only data. Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Jinfeng Xu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
AAAI | 4 |
| 2025 | SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point CloudsabstractRecent advancements in deep learning have greatly enhanced 3D object recognition, but most models are limited to closed-set scenarios, unable to handle unknown samples in real-world applications. Open-set recognition (OSR) addresses this limitation by enabling models to both classify known classes and identify novel classes. However, current OSR methods rely on global features to differentiate known and unknown classes, treating the entire object uniformly and overlooking the varying semantic importance of its different parts. To address this gap, we propose Salience-Aware Structured Separation (SASep), which includes (i) a tunable semantic decomposition (TSD) module to semantically decompose objects into important and unimportant parts, (ii) a geometric synthesis strategy (GSS) to generate pseudo-unknown objects by combining these unimportant parts, and (iii) a synth-aided margin separation (SMS) module to enhance feature-level separation by expanding the feature distributions between classes. Together, these components improve both geometric and feature representations, enhancing the model’s ability to effectively distinguish known and unknown classes. Experimental results show that SASep achieves superior performance in 3D OSR, outperforming existing state-of-the-art methods. The codes are available at https://github.com/JinfengX/SASep. Jinfeng Xu 0002, Xianzhi Li 0001, Xu Han 0016, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
CVPR | 5 |
| 2025 | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play DeformationabstractGenerating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images, and use the Large Reconstruction Model (LRM) to create the final meshes. However, the multiview images exhibit local inconsistencies, and the meshes often lack fidelity to the input image or look blurry. We propose Fancy123, featuring two enhancement modules and an unprojection operation to address the above three issues, respectively. The appearance enhancement module deforms the 2D multiview images to realign misaligned pixels for better multiview consistency. The fidelity enhancement module deforms the 3D mesh to match the input image. The unprojection of the input image and deformed multiview images onto LRM’s generated mesh ensures high clarity, discarding LRM’s predicted blurry-looking mesh colors. Extensive qualitative and quantitative experiments verify Fancy123’s SoTA performance with significant improvement. Also, the two enhancement modules are plug-and-play and work at inference time, allowing seamless integration into various existing single-image-to-3D methods. Project page: https://github.com/YuQiao0303/Fancy123. Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Long Hu, Yixue Hao, Min Chen 0003 |
CVPR | 1 |
| 2025 | PointDreamer: Zero-Shot 3D Textured Mesh Reconstruction From Colored Point CloudabstractFaithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) offers inherent advantages in comprehensively and precisely replicating the target object's 360$^{\circ }$∘ characteristics. While most existing colored-PC-to-mesh methods suffer from blurry textures or require hard-to-acquire 3D training data, we propose PointDreamer, a novel framework that harnesses 2D diffusion prior for superior texture quality. Crucially, unlike prior 2D-diffusion-for-3D works driven by text or image inputs, PointDreamer successfully adapts 2D diffusion models to 3D point cloud data by a novel project-inpaint-unproject pipeline. Specifically, it first projects the point cloud into sparse 2D images and then performs diffusion-based inpainting. After that, diverging from most existing 3D reconstruction or generation approaches that predict texture in 3D/UV space thus often yielding blurry texture, PointDreamer achieves high-quality texture by directly unprojecting the inpainted 2D images to the 3D mesh. Furthermore, we identify for the first time a typical kind of unprojection artifact appearing in occlusion borders, which is common in other multiview-image-to-3D pipelines but less-explored. To address this, we propose a novel solution named the Non-Border-First (NBF) unprojection strategy. Extensive qualitative and quantitative experiments on various synthetic and real-scanned datasets demonstrate that PointDreamer, though zero-shot, exhibits SoTA performance (30% improvement on LPIPS score from 0.118 to 0.068), and is robust to noisy, sparse, or even incomplete input data. Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Jinfeng Xu 0002, Long Hu, Min Chen 0003 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | JIMR: Joint Semantic and Geometry Learning for Point Scene Instance Mesh ReconstructionabstractPoint scene instance mesh reconstruction is a challenging task since it requires both scene-level instance segmentation and instance-level mesh reconstruction from partial observations simultaneously. Previous works either adopt a detection backbone or a segmentation one, and then directly employ a mesh reconstruction network to produce complete meshes from incomplete instance point clouds. To further boost the mesh reconstruction quality with both local details and global smoothness, in this work, we propose JIMR, a joint framework with two cascaded stages for semantic and geometry understanding. In the first stage, we propose to perform both instance segmentation and object detection simultaneously. By making both tasks promote each other, this design facilitates subsequent mesh reconstruction by providing more precisely-segmented instance points and better alignment benefiting from predicted complete bounding boxes. In the second stage, we propose a complete-then-reconstruct procedure, where the completion module explicitly disentangles completion from reconstruction, and enables the usage of pre-trained weights of existing powerful completion and reconstruction networks. Moreover, we propose a comprehensive confidence score to filter proposals considering the quality of instance segmentation, bounding box detection, semantic classification, and mesh reconstruction at the same time. Experiments show that our proposed JIMR outperforms state-of-the-art methods regarding instance reconstruction qualitatively and quantitatively. Qiao Yu 0002, Xianzhi Li 0001, Jinfeng Xu 0002, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D PriorsabstractLarge 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (3D-LLMs) also integrate point clouds into LLMs. However, directly aligning point clouds with LLM requires expensive training costs, typically in hundreds of GPU-hours on A100, which hinders the development of 3D-LLMs. In this paper, we introduce MiniGPT-3D, an efficient and powerful 3D-LLM that achieves multiple SOTA results while training for only 27 hours on one RTX 3090. Specifically, we propose to align 3D point clouds with LLMs using 2D priors from 2D-LLMs, which can leverage the similarity between 2D and 3D visual information. We introduce a novel four-stage training strategy for modality alignment in a cascaded way, and a mixture of query experts module to adaptively aggregate features with high efficiency. Moreover, we utilize parameter-efficient fine-tuning methods LoRA and Norm fine-tuning, resulting in only 47.8M learnable parameters, which is up to 260x fewer than existing methods. Extensive experiments show that MiniGPT-3D achieves SOTA on 3D object classification and captioning tasks, with significantly cheaper training costs. Notably, MiniGPT-3D gains an 8.12 increase on GPT-4 evaluation score for the challenging object captioning task compared to ShapeLLM-13B, while the latter costs 160 total GPU-hours on 8 A800. We are the first to explore the efficient 3D-LLM, offering new insights to the community. Code and weights are available at https://github.com/TangYuan96/MiniGPT-3D. Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
ACM Multimedia | 4 |
| 2024 | Point-LGMask: Local and Global Contexts Embedding for Point Cloud Pre-Training With Multi-Ratio MaskingabstractSelf-supervised learning has achieved great success in both natural language processing and 2D vision, where masked modeling is a quite popular pre-training scheme. However, extending masking to 3D point cloud understanding that combines local and global features poses a new challenge. In our work, we present Point-LGMask, a novel method to embed both local and global contexts with multi-ratio masking, which is quite effective for self-supervised feature learning of point clouds but is unfortunately ignored by existing pre-training works. Specifically, to avoid fitting to a fixed masking ratio, we first propose multi-ratio masking, which prompts the encoder to fully explore representative features thanks to tasks of different difficulties. Next, to encourage the embedding of both local and global features, we formulate a compound loss, which consists of (i) a global representation contrastive loss to encourage the cluster assignments of the masked point clouds to be consistent to that of the completed input, and (ii) a local point cloud prediction loss to encourage accurate prediction of masked points. Equipped with our Point-LGMask, we show that our learned representations transfer well to various downstream tasks, including few-shot classification, shape classification, object part segmentation, as well as real-world scene-based 3D object detection and 3D semantic segmentation. Particularly, our model largely advances existing pre-training methods on the difficult few-shot classification task using the real-captured ScanObjectNN dataset by surpassing over 4% to the second-best method. Also, our Point-LGMask achieves 0.4%$AP_{25}$and 0.8%$AP_{50}$gains on 3D object detection task over the second-best method. 0.4% mAcc and 0.5% mIoU. Codes have been released athttps://github.com/TangYuan96/Point-LGMask. Xianzhi Li 0001, Jinfeng Xu 0002, Qiao Yu 0002, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Multim. | 4 |
| 2023 | CasFusionNet: A Cascaded Network for Point Cloud Semantic Scene Completion by Dense Feature FusionabstractSemantic scene completion (SSC) aims to complete a partial 3D scene and predict its semantics simultaneously. Most existing works adopt the voxel representations, thus suffering from the growth of memory and computation cost as the voxel resolution increases. Though a few works attempt to solve SSC from the perspective of 3D point clouds, they have not fully exploited the correlation and complementarity between the two tasks of scene completion and semantic segmentation. In our work, we present CasFusionNet, a novel cascaded network for point cloud semantic scene completion by dense feature fusion. Specifically, we design (i) a global completion module (GCM) to produce an upsampled and completed but coarse point set, (ii) a semantic segmentation module (SSM) to predict the per-point semantic labels of the completed points generated by GCM, and (iii) a local refinement module (LRM) to further refine the coarse completed points and the associated labels from a local perspective. We organize the above three modules via dense feature fusion in each level, and cascade a total of four levels, where we also employ feature fusion between each level for sufficient information usage. Both quantitative and qualitative results on our compiled two point-based datasets validate the effectiveness and superiority of our CasFusionNet compared to state-of-the-art methods in terms of both scene completion and semantic segmentation. The codes and datasets are available at: https://github.com/JinfengX/CasFusionNet. Jinfeng Xu 0002, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
AAAI | 4 |
| 2023 | Intelligent Fabric Enabled 6G Semantic Communication System for In-Cabin ScenariosabstractWith the large-scale commercialization of 5G, the global industry has started the exploration of the next generation mobile communication technology (6G). From mobile Internet, to IoT, and then to the smart connection of everything, 6G will transform from 5G’s service objects of people and things to the intelligent networking of agent that supports human–machine–object. 6G networks should have the characteristics of ubiquitous intelligence and ubiquitous perception, which poses challenges for 6G network construction. Therefore, we propose a 6G Semantic Communication Scheme based on Intelligent Fabrics for transportation in-cabin scenarios (6GSCS-IF), which can provide senseless intelligent interaction in transportation in-cabin environment through widely and flexibly deployed intelligent fabrics, demonstrating the superiority of intelligent fabrics in realizing human–machine–object intelligent sensory interaction. Then, we propose a Deep Learning-based Semantic Communication Model for Time-series data (DL-SCMT), and use deep learning for semantic sensing and information extraction to build an end-to-end semantic communication system. The experimental results show that the semantic communication services provided by this model can achieve better signal reconstruction and higher-order intelligent services compared with traditional communication methods. Qiao Yu 0002, Di Wu 0001, Chong Hou, Guangming Tao, Min Chen 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | GNN-Based Depression Recognition Using Spatio-Temporal Information: A fNIRS StudyabstractIn recent years, depression has become an increasingly serious problem globally. Previous studies of automatic depression recognition based on functional near-Infrared spectroscopy (fNIRS) or other brain imaging techniques have shown potential to serve as auxiliary diagnosis methods that provide assistance to medical professionals. Recently, some studies have found that, besides directly using the data themselves (temporal data), the use of functional connectivity among channels (spatial data) also can be effective. In this paper, we propose a method based on Graph Neural Network (GNN) that combines both temporal and spatial features of fNIRS data for automatic depression recognition. Specifically, fNIRS data of 96 subjects were collected and pre-processed. Basic statistical metrics of each channel were extracted as temporal features, and channel connectivity (coherence and correlation) were calculated as spatial features. Point-biserial analysis was conducted on these features and depression labels as a data-driven motivation. For classification, we considered data of each subject as a graph, with temporal features as node features and spatial features as edge weights. The graphs were fed into GNNs for training and testing. Experimental results showed that our GNN-based methods realized the best depression recognition performance compared with classical machine-learning methods regarding accuracy, F1 score, and precision, especially in F1 score for over 10%. Qiao Yu 0002, Rui Wang 0077, Jia Liu 0009, Long Hu, Min Chen 0003, Zhongchun Liu |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Medical-Level Suicide Risk Analysis: A Novel Standard and Evaluation ModelabstractThe frequent occurrence of suicides in modern society constitutes a serious public health issue. While the motives, methods, and consequences of suicide are quite complicated, if people at risk of suicide can be identified and intervened in time, the loss of life can be reduced. Through analyses based on combining a large number of suicide texts and professional medical literature, a dictionary of potential suicide risk impact factors has been established in this article. Based on this dictionary, a novel medical-level suicide risk standard is proposed to monitor suicide risk from point-to-surface under the timeline baseline. In order to solve the problem of insufficient Chinese suicide data sets, the manually assisted method based on knowledge perception is adopted to annotate the data set with corresponding to risk level. At the same time, a Bert evaluation model based on knowledge perception was established for the classification of risk level. The experimental results showed that proposed method has a 56% recognition accuracy in the prediction of 10-Label suicide risk level proposed in this article, and the classification performance is better than traditional machine learning algorithms. Therefore, the results showed that the classification standard and evaluation model can be effectively used for the identification and early warning of suicide risk, which can discover high suicide risk groups to reduce the occurrence of suicide. It is of great significance to people’s emotion care monitoring. Rui Wang 0077, Bing Xiang Yang, Yujun Ma, Qiao Yu 0002, Xiaofen Zong, Simeng Ma, Long Hu, Kai Hwang 0001, Zhongchun Liu |
IEEE Internet Things J. | 5 |
| 2021 | Depression Analysis and Recognition Based on Functional Near-Infrared SpectroscopyabstractDepression is the result of a complex interaction of social, psychological and physiological elements. Research into the brain disorders of patients suffering from depression can help doctors to understand the pathogenesis of depression and facilitate its diagnosis and treatment. Functional near-infrared spectroscopy (fNIRS) is a non-invasive approach to the detection of brain functions and activities. In this paper, a comprehensive fNIRS-based depression-processing architecture, including the layers of source, feature and model, is first established to guide the deep modeling for fNIRS. In view of the complexity of depression, we propose a methodology in the time and frequency domains for feature extraction and deep neural networks for depression recognition combined with current research. It is found that compared to non-depression people, patients with depression have a weaker encephalic area connectivity and lower level of activation in the prefrontal lobe during brain activity. Finally, based on raw data, manual features and channel correlations, the AlexNet model shows the best performance, especially in terms of the correlation features and presents an accuracy rate of 0.90 and a precision rate of 0.91, which is higher than ResNet18 and machine-learning algorithms on other data. Therefore, the correlation of brain regions can effectively recognize depression (from cases of non-depression), making it significant for the recognition of brain functions in the clinical diagnosis and treatment of depression. Rui Wang 0077, Yixue Hao, Qiao Yu 0002, Min Chen 0003, Iztok Humar, Giancarlo Fortino |
IEEE J. Biomed. Health Informatics | 3 |