Fukun Yin

dblp:272/0842 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-2623-1619ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 10 since 2021
YearPublicationVenuePosition
2026 A Unified multi-modality conditional latent diffusion model for point cloud generation
Yihang Yang, Zibo Zhao 0001, Fukun Yin, Wen Liu 0003, Yuhan Ding, Biao Jiang, Gang Yu 0002, Tao Chen 0003
Pattern Recognit.3
2026 Rendered 2D Semantic and Generative Priors Guided 3D Multi-Object Grounding
abstract
3D multi-object visual grounding aims to identify and localize all objects in a 3D scene that correspond to a given text description. Unlike traditional single-object grounding, this task presents additional challenges as point clouds inherently lack fine-grained details, making it difficult to capture subtle object features and contextual information. Moreover, textual descriptions are inherently limited in perceiving and understanding complex 3D environments, especially in scenarios with high object similarity or intricate spatial arrangements. To tackle the above challenges, we propose SGMG, a Rendered 2D Semantic and Generative priors guided 3D Multi-object Grounding Framework. The SGMG framework introduces two key innovations that work cohesively to enhance grounding accuracy. First, a Generative-Assistant(GA) Module leverages the capabilities of a generative model to provide enriched scene prior information and capture the fine-grained scene details. Second, the Semantic-Augment Fusion(SAF) Module is designed to improve the representation of text features from the vision features, thereby boosting the accuracy of multimodal information interactions. Furthermore, we introduce a multi-level fusion mechanism, ensuring that semantic and spatial relationships between objects are preserved and effectively leveraged during the grounding process. Experimental results demonstrate that SGMG achieves state-of-the-art performance in multi-object 3D grounding and competitive results in traditional single-object tasks, highlighting its effectiveness in diverse scenarios.
Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Yihang Yang, Fukun Yin, Tao Chen 0003
IEEE Trans. Multim.5
2025 A Multimodal LLM for Chart Understanding and Generation
abstract
Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for specific domain data, particularly when it comes to interpreting chart figures. This is mainly due to the lack of relevant multi-modal instruction tuning datasets. In this article, we create a high-quality instruction-tuning dataset leveraging GPT-4. We develop a multi-step data generation process in which different steps are responsible for generating tabular data, creating chart figures, and designing instruction tuning data separately. Our method’s flexibility enables us to generate diverse, high-quality instruction-tuning data consistently and efficiently while maintaining a low resource expenditure. Additionally, it allows us to incorporate a wider variety of chart and task types not yet featured in existing datasets. Next, we introduce ChartLlama, a multi-modal large language model that we’ve trained using our created dataset. ChartLlama outperforms all prior methods in ChartQA, Chart-to-text, and Chart-extraction evaluation benchmarks. Additionally, ChartLlama significantly improves upon the baseline in our specially compiled chart dataset, which includes new chart and task types. The results of ChartLlama confirm the value and huge potential of our proposed data generation method in enhancing chart comprehension.
Yucheng Han, Chi Zhang 0007, Xin Chen 0040, Fukun Yin, Xu Yang 0021, Zhibin Wang 0004, Gang Yu 0002, Hanwang Zhang
IJCNN4
2025 Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAE
abstract
As Artificial Intelligence Generated Content (AIGC) advances, a variety of methods have been developed to generate text, images, videos, and 3D shapes from single or multimodal inputs, contributing efforts to emulate human-like cognitive content creation. However, generating realistic large-scale scenes from a single input presents a challenge due to the complexities involved in ensuring consistency across extrapolated views generated by models. Benefiting from recent video generation models and implicit neural representations, we propose Scene123, a 3D scene generation model, which combines a video generation framework to ensure realism and diversity with implicit neural fields integrated with Masked Autoencoders (MAE) to effectively ensure the consistency of unseen areas across views. Specifically, the input image (or a text-generated image) is first warped to simulate adjacent views, with the invisible regions filled using the consistency-enhanced MAE model. Nonetheless, the synthesized images often exhibit inconsistencies in viewpoint alignment, thus we utilize the produced views to optimize a neural radiance field, enhancing geometric consistency. Moreover, to further enhance the details and texture fidelity of generated views, we employ a GAN-based Loss against images derived from the input image through the video generation model. Extensive experiments demonstrate that our method can generate realistic and consistent scenes from a single prompt. Both qualitative and quantitative results indicate that our approach surpasses existing state-of-the-art methods.
Fukun Yin, Jiayuan Fan 0001, Wanzhang Li, Xin Chen 0040, Gang Yu 0002
ACM Multimedia2
2025 OmniSVG: A Unified Scalable Vector Graphics Generation Model
abstract
Scalable Vector Graphics (SVG) is an important image format widely adopted in graphic design because of their resolution independence and editability. The study of generating high-quality SVG has continuously drawn attention from both designers and researchers in the AIGC community. However, existing methods either produces unstructured outputs with huge computational cost or is limited to generating monochrome icons of over-simplified structures. To produce high-quality and complex SVG, we propose OmniSVG, a unified framework that leverages pre-trained Vision-Language Models (VLMs) for end-to-end multimodal SVG generation. By parameterizing SVG commands and coordinates into discrete tokens, OmniSVG decouples structural logic from low-level geometry for efficient training while maintaining the expressiveness of complex SVG structure. To further advance the development of SVG synthesis, we introduce MMSVG-2M, a multimodal dataset with two million richly annotated SVG assets, along with a standardized evaluation protocol for conditional SVG generation tasks. Extensive experiments show that OmniSVG outperforms existing methods and demonstrates its potential for integration into professional SVG design workflows.
Sijin Chen, Xianfang Zeng, Fukun Yin, Gang Yu 0002, Xingjun Ma, Yu-Gang Jiang 0001
NeurIPS5
2025 Corrections to "DCNet: Large-Scale Point Cloud Semantic Segmentation With Discriminative and Efficient Feature Aggregation"
abstract
Presents corrections to the paper, “DCNet: Large-Scale Point Cloud Semantic Segmentation With Discriminative and Efficient Feature Aggregation”.
Fukun Yin, Tao Chen 0003, Guozhong Luo, Gang Yu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2025 WI3D: Weakly Incremental 3D Detection via Vision Foundation Models
abstract
Class-incremental 3D object detection demands a 3D detector tolocateandrecognizenovel categories in a stream fashion while preserving its base detection ability. However, existing methods require delicate 3D annotations for learning novel categories, resulting in significant labeling costs. To this end, we explore a label-efficient approach calledWeaklyIncremental3DDetection (WI3D), which teaches a 3D detector to learn incrementally with off-the-shelf vision foundation models. We propose a novel dual-teaching framework incorporating both intra-modal and inter-modal knowledge from pseudo labels and feature space. Specifically, our framework features a class-agnostic pseudo-label refinement module, designed for the generation of high-quality 3D pseudo labels. This module is built on a lightweight transformer that models the spatial relationships between pseudo labels and their interactions with rich contextual information in point clouds. Additionally, we introduce a cross-modal knowledge transfer module to enhance the representation learning of novel classes, along with a reweighting knowledge distillation strategy that dynamically assesses and distills knowledge from previously learned categories. Extensive experiments show that our approach can efficiently learn novel concepts while preserving knowledge of base classes in WI3D scenarios, and surpass baseline approaches on both SUN-RGBD and ScanNet.
Mingsheng Li, Sijin Chen, Shengji Tang, Hongyuan Zhu 0002, Yanyan Fang, Xin Chen 0040, Zhuoyuan Li 0006, Fukun Yin, Tao Chen 0003
IEEE Trans. Multim.8
2025 ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language Model
abstract
The advent of large language models, which enable flexibility through instruction-driven approaches, has revolutionized many traditional generative tasks, but large models for 3D data, particularly in comprehensively handling 3D shapes with other modalities, are still under-explored. By achieving instruction-based shape generation, versatile multi-modal generative shape models can significantly benefit various fields, such as 3D virtual construction and network-aided design. In this article, we present ShapeGPT, a shape-included multi-modal framework to leverage strong pre-trained language models to address multiple shape-relevant tasks. Specifically, ShapeGPT employs a “word-sentence-paragraph” framework to discretize continuous shapes into shape words, further assembles these words into shape sentences, and integrates shape with instructional text for multi-modal paragraphs. To learn this shape-language model, we use a three-stage training scheme, including shape representation, multi-modal alignment, and instruction-based generation, to align shape-language codebooks and learn the intricate correlations among these modalities. Extensive experiments demonstrate that ShapeGPT achieves comparable performance across shape-relevant tasks, including text-to-shape, shape-to-text, shape completion, and shape editing.
Fukun Yin, Xin Chen 0040, Chi Zhang 0007, Biao Jiang, Zibo Zhao 0001, Wen Liu 0003, Gang Yu 0002, Tao Chen 0003
IEEE Trans. Multim.1
2024 PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural Representation
abstract
Recent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional sampling cannot cope with the cubically growing sampling space. To alleviate the dependence on filling the sampling space, we explore using multi-modal priors to assist individual points to obtain more global semantic information and propose a priorrich multi-modal implicit neural representation network, Pm-INR, for the outdoor unbounded large-scale scene. The core of our method is multi-modal prior extraction and crossmodal prior fusion modules. The former encodes codebooks from different modality inputs and extracts valuable priors, while the latter fuses priors to maintain view consistency and preserve unique features among multi-modal priors. Finally, feature-rich cross-modal priors are injected into the sampling regions to allow each region to perceive global information without filling the sampling space. Extensive experiments have demonstrated the effectiveness and robustness of our method for outdoor unbounded large-scale scene novel view synthesis, which outperforms state-of-the-art methods in terms of PSNR, SSIM, and LPIPS.
Fukun Yin, Wen Liu 0003, Jiayuan Fan 0001, Xin Chen 0040, Gang Yu 0002, Tao Chen 0003
AAAI2
2024 MotionChain: Conversational Motion Controllers via Multimodal Prompts
Biao Jiang, Xin Chen 0040, Chi Zhang 0007, Fukun Yin, Zhuoyuan Li 0006, Gang Yu 0002, Jiayuan Fan 0001
ECCV (26)4
2024 M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions
Mingsheng Li, Xin Chen 0040, Chi Zhang 0007, Sijin Chen, Hongyuan Zhu 0002, Fukun Yin, Zhuoyuan Li 0006, Gang Yu 0002, Tao Chen 0003
ECCV (58)6
2024 G-Former: A Grouping Transformer for Weakly Supervised Point Cloud Segmentation
abstract
Recent advancements in weakly supervised point cloud semantic segmentation have diminished the reliance on extensive annotations, thereby enhancing the efficacy of understanding the real-world environment. However, existing approaches, such as PSD [25], SQN [5] and OTOC [10], often overlook the valuable global class-related prior knowledge present in point clouds beyond the scope of labels. To fully leverage this prior knowledge, which suggests that points of the same class should be close in feature space and each class should have a representative feature, we propose G-Former, a grouping transformer model. G-Former incorporates the idea of clustering into the overall model by defining clusters aligned to classes and assigning learnable tensors as cluster centers. Points are then grouped into these clusters based on the similarity of their features to the cluster centers. The core components of G-Former include a Hierarchy Cluster Structure (HCS) and a Grouping Module (GM). The former consists of two sets of clusters, one for classes while the other serves as a middle layer to help class clusters handle large-scale point features. The latter facilitates grouping the point cloud into different clusters. With the help of the grouping transformer model, G-Former further proposes a series of cluster center constraints to augment inter-class distances and diminish intra-class distances to enhance the discriminability of points. Experimental results on ScanNet v2 and S3DIS datasets demonstrate that G-Former outperforms previous methods with limited labels (0.1% or 1%) by a significant margin and is even comparable to fully supervised methods.
Zehan Huang, Fukun Yin, Jiayuan Fan 0001, Xin Chen 0040, Hongyuan Zhu 0002, Bin Wang 0008, Tao Chen 0003
IJCNN2
2024 CSD3D: Cross-Scale Distillation via Dual-Consistency Learning for Semi-Supervised 3D Object Detection
abstract
Semi-supervised 3D object detection has gained significant attention due to its potential to mitigate the heavy reliance on extensive annotations in traditional 3D object detection methodologies. Most existing approaches leverage the teacher’s predictions to guide and refine the student’s predictions while discarding low-confidence predictions using a fixed threshold. However, the presence of imbalanced variance in object scale poses challenges as different objects often exhibit varying levels of detection difficulty. Current methods relying on pseudo labels struggle to comprehensively capture information pertaining to objects across diverse scales. To address these challenges, we propose CSD3D, a cross-scale distillation approach via dual-consistency learning. CSD3D encompasses cross-scale distillation between the teacher and student as well as within the student itself, thereby enhancing the algorithm’s resilience to scale variance. Moreover, by adopting a dual-consistency learning paradigm that incorporates supervision at both feature and prediction levels, our approach provides comprehensive guidance to the student model. This integration of dual-consistency learning within cross-scale conditions is conducive to comprehending cross-scale object features and maintaining scale-consistent predictions. Rigorous experiments performed on the ScanNet and SUN RGB-D benchmarks reveal that CSD3D attains state-of-the-art performance. By utilizing a mere 10% of labeled data on ScanNet, we observe absolute improvements of 3.8 and 3.5 in [email protected] and [email protected], respectively.
Sikai Wu, Fukun Yin, Hancheng Ye, Tao Chen 0003
IJCNN2
2024 MeshXL: Neural Coordinate Field for Generative 3D Foundation Models
abstract
The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately, with a pre-defined ordering strategy, 3D meshes can be represented as sequences, and the generation process can be seamlessly treated as an auto-regressive problem. In this paper, we validate Neural Coordinate Field (NeurCF), an explicit coordinate representation with implicit neural embeddings, is a simple-yet-effective representation for large-scale sequential mesh modeling. After that, we present MeshXL, a family of generative pre-trained auto-regressive models that addresses 3D mesh generation with modern large language model approaches. Extensive experiments show that MeshXL is able to generate high-quality 3D meshes, and can also serve as foundation models for various down-stream applications.
Sijin Chen, Xin Chen 0040, Anqi Pang, Xianfang Zeng, Yijun Fu, Fukun Yin, Billzb Wang, Jingyi Yu 0001, Gang Yu 0002, Tao Chen 0003
NeurIPS7
2024 Instruct Pix-to-3D: Instructional 3D object generation from a single image
Wen Liu 0003, Wanzhang Li, Zibo Zhao 0001, Fukun Yin, Xin Chen 0018, Lei Zhao 0035, Tao Chen 0003
Neurocomputing5
2023 A Large-Scale Outdoor Multi-modal Dataset and Benchmark for Novel View Synthesis and Implicit Scene Reconstruction
abstract
Neural Radiance Fields (NeRF) [24] has achieved impressive results in single object scene reconstruction and novel view synthesis, as demonstrated on many single modality and single object focused indoor scene datasets like DTU [14], BMVS [42], and NeRF Synthetic [24]. However, the study of NeRF on large-scale outdoor scene reconstruction is still limited, as there is no unified outdoor scene dataset for large-scale NeRF evaluation due to expensive data acquisition and calibration costs. In this work, we propose a large-scale outdoor multi-modal dataset, OMMO dataset, containing complex objects and scenes with calibrated images, point clouds and prompt annotations. A new benchmark for several outdoor NeRF-based tasks is established, such as novel view synthesis, diverse 3D representation, and multi-modal NeRF. To create the dataset, we capture and collect a large number of real fly-view videos and select high-quality and high-resolution clips from them. Then we design a quality review module to refine images, remove low-quality frames and fail-to-calibrate scenes through a learning-based automatic evaluation plus manual review. Finally, volunteers are employed to label and review the prompt annotation for each scene and keyframe. Compared with existing NeRF datasets, our dataset contains abundant real-world urban and natural scenes with various scales, camera trajectories, and lighting conditions. Experiments show that our dataset can benchmark most state-of-the-art NeRF methods on different tasks. The dataset can be found at the following link: https://ommo.luchongshan.com/.
Chongshan Lu, Fukun Yin, Xin Chen 0040, Wen Liu 0003, Tao Chen 0003, Gang Yu 0002, Jiayuan Fan 0001
ICCV2
2023 PDF: Point Diffusion Implicit Function for Large-scale Scene Neural Representation
abstract
Recent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge for unbounded large-scale outdoor scenes. To alleviate the dilemma of using individual points to perceive the entire colossal space, we explore learning the surface distribution of the scene to provide structural priors and reduce the samplable space and propose a Point Diffusion implicit Function, PDF, for large-scale scene neural representation. The core of our method is a large-scale point cloud super-resolution diffusion module that enhances the sparse point cloud reconstructed from several training images into a dense point cloud as an explicit prior. Then in the rendering stage, only sampling points with prior points within the sampling radius are retained. That is, the sampling space is reduced from the unbounded space to the scene surface. Meanwhile, to fill in the background of the scene that cannot be provided by point clouds, the region sampling based on Mip-NeRF 360 is employed to model the background representation. Expensive experiments have demonstrated the effectiveness of our method for large-scale scene novel view synthesis, which outperforms relevant state-of-the-art baselines.
Yuhan Ding, Fukun Yin, Jiayuan Fan 0001, Xin Chen 0040, Wen Liu 0003, Chongshan Lu, Gang Yu 0002, Tao Chen 0003
NeurIPS2
2023 DCNet: Large-Scale Point Cloud Semantic Segmentation With Discriminative and Efficient Feature Aggregation
abstract
The point cloud feature aggregation, which learns discriminative features from the disordered points, plays a key role for large-scale point cloud semantic segmentation. Most previous aggregation methods are based on sampling a representative point subset, i.e., by a carefully designed point density metric, facing expensive computation cost especially for large-scale point clouds. Even though speeding up the point sampling process is studied by several recent works, but the component points in the sampled subset are uncertain and may change randomly, thus leading to corrupted geometric structure and discarded edges for representing an object. Therefore, we propose the DCNet, which consists of a fast point random sampling based encoder-decoder structure and several fully connected layers for semantic segmentation. To overcome the key feature loss caused by random down-sampling, the DCNet develops two novel local feature aggregation schemes: Double attention and Consistent constraints, to learn features that are discriminative for the challenging scenarios as above. The former considers both topological and semantic similarity of neighboring points to generate attention features for discriminating classes with similar geometric structures. The latter develops class-consistent constraints between adjacent layers in the decoder stage, to guide each point to aggregate with high-level semantic features of points belonging to the same class from the previous layer, which is beneficial for distinguishing neighboring points of the same class on the boundary. We conduct experiments and compare the proposed DCNet with existing methods on two benchmarks S3DIS and Semantic3D. Experiments show that the mean Intersection-over-Union (mIoU) of our method outperforms state-of-the-art methods by 2-3%, based on the same fast random sampling, and is also comparable to latest sampling-slower but accuracy-higher methods. That is, our method achieves the optimal speed-accuracy trade-off in the field of point cloud segmentation.
Fukun Yin, Tao Chen 0003, Guozhong Luo, Gang Yu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2022 Coordinates Are NOT Lonely - Codebook Prior Helps Implicit Neural 3D representations
abstract
Implicit neural 3D representation has achieved impressive results in surface or scene reconstruction and novel view synthesis, which typically uses the coordinate-based multi-layer perceptrons (MLPs) to learn a continuous scene representation. However, existing approaches, such as Neural Radiance Field (NeRF) and its variants, usually require dense input views (i.e. 50-150) to obtain decent results. To relive the over-dependence on massive calibrated images and enrich the coordinate-based feature representation, we explore injecting the prior information into the coordinate-based network and introduce a novel coordinate-based model, CoCo-INR, for implicit neural 3D representation. The cores of our method are two attention modules: codebook attention and coordinate attention. The former extracts the useful prototypes containing rich geometry and appearance information from the prior codebook, and the latter propagates such prior information into each coordinate and enriches its feature representation for a scene or object surface. With the help of the prior information, our method can render 3D views with more photo-realistic appearance and geometries than the current methods using fewer calibrated images available. Experiments on various scene reconstruction datasets, including DTU and BlendedMVS, and the full 3D head reconstruction dataset, H3DS, demonstrate the robustness under fewer input views and fine detail-preserving capability of our proposed method.
Fukun Yin, Wen Liu 0003, Tao Chen 0003, Gang Yu 0002
NeurIPS1
2020 Accurate Estimation of Body Height From a Single Depth Image via a Four-Stage Developing Network
abstract
Non-contact measurement of human body height can be very difficult under some circumstances.In this paper we address the problem of accurately estimating the height of a person with arbitrary postures from a single depth image. By introducing a novel part-based intermediate representation plus a four-stage increasingly complex deep neural network, we manage to achieve significantly higher accuracy than previous methods. We first describe the human body in the form of a segmentation of human torso as four nearly rigid parts and then predict their lengths respectively by 3 CNNs. Instead of directly adding the lengths of these parts together, we further construct another independent developing CNN that combines the intermediate representation, part lengths and depth information together to finally predict the body height results.Here we develop an increasingly complex network architecture and adopt a hybrid pooling to optimize training process. To the best of our knowledge, this is the first method that estimates height only from a single depth image. In experiments our average accuracy reaches at 99.1% for people in various positions and postures.
Fukun Yin, Shizhe Zhou
CVPR1