EDBT 2026 Demo / reviewers in the wild / expert
Yanping Fu
dblp:194/0483
· DBLP profile ↗
35ranked-venue papers
18as first author
27since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 10 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 9 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Texture and Geometry Optimization for 3D Reconstruction
Yanping Fu, Hongjing Zhang, Shaojie Zhang 0002, Dengdi Sun, Haifeng Zhao 0001 |
CGI (1) | 1 |
| 2025 | Sparse-View X-ray 3D Reconstruction using Hybrid Representation Neural Attenuation FieldsabstractX-ray 3D reconstruction has achieved superior performance in medical imaging with traditional and deep learning methods. However, when sparse-view X-ray projections are used to minimize patient exposure to radiation, these methods tend to overfit and produce blurring. To overcome this problem, we propose a novel hybrid feature representation neural attenuation field framework for sparse-view X-ray 3D reconstruction. First, we integrate tri-plane features with hash coding features as the network input, enabling the capture of intricate local details and high-frequency information. Second, we enhance the modeling of radiation attenuation across different organs by designing a specialized attenuation weight estimation network. This network enables the attenuation field estimation network to more accurately focus on the varying attenuation rates of different tissues. Third, we introduce a new multi-skip strategy, which uses the skip connection strategy for each layer of MLPs to the attenuation value and weight prediction network, markedly enhancing the performance of our method. Experiments on public datasets demonstrate the superiority of our proposed method over state-of-the-art methods. Yanping Fu, Hao Geng, Zhuangzhuang Zhao, Shaojie Zhang 0002, Haifeng Zhao 0001 |
ICASSP | 1 |
| 2025 | Single-View Clothed Human Reconstruction using Symmetric FeatureabstractSingle-view clothed human reconstruction has emerged as a prominent research focus, yet reconstructing the occluded backside of the human body remains a significant challenge due to viewpoint limitations and occlusion. In this paper, we propose a novel single-view clothed human reconstruction method that leverages non-rigid symmetric features guidance to effectively mitigate these challenges. Firstly, we utilize the natural symmetry of the human body to construct the non-rigid symmetry through the parametric human body model SMPL-X. Secondly, these symmetric-driven features are then used to enhance the pixel-aligned features, enabling a more precise and complete reconstruction. Finally, we introduce an innovative occupancy query selector that intelligently identifies and prioritizes the most relevant features during the query stage to further improve the quality of reconstruction. Experimental results demonstrate that the proposed method can achieve better reconstruction results, especially in occluded areas. Yanping Fu, Zhuangzhuang Zhao, Hao Geng, Haifeng Zhao 0001 |
ICASSP | 1 |
| 2025 | RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General GraspingabstractGeneral robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-world effectiveness. To address this limitation, we build a large-scale grasping-oriented affordance segmentation benchmark with human-like instructions, named RAGNet. It contains 273k images, 180 categories, and 26k reasoning instructions. The images cover diverse embodied data domains, such as wild, robot, ego-centric, and even simulation data. They are carefully annotated with an affordance map, while the difficulty of language instructions is largely increased by removing their category name and only providing functional descriptions. Furthermore, we propose a comprehensive affordance-based grasping framework, named AffordanceNet, which consists of a VLM pre-trained on our massive affordance data and a grasping network that conditions an affordance map to grasp the target. Extensive experiments on affordance segmentation benchmarks and real-robot manipulation tasks show that our model has a powerful open-world generalization ability. Our data and code is available at https://github.com/wudongming97/AffordanceNet. Dongming Wu 0005, Yanping Fu, Saike Huang, Yingfei Liu, Fan Jia 0006, Nian Liu 0002, Tiancai Wang, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jianbing Shen |
ICCV | 2 |
| 2025 | PFGM-IQA: CT Image Quality Assessment using Poisson Flow Generative ModelsabstractImage Quality Assessment (IQA) is essential for optimizing radiation dose in Computed Tomography (CT) while ensuring radiologists achieve the highest diagnostic accuracy. To advance research in this field, we propose a novel algorithm for CT image quality assessment without the need for reference images. To overcome the challenge of no-reference IQA in CT scans, we employ a novel approach using Poisson Flow Generative models (PFGM) to generate pseudo-reference images for low-dose CT scans. These pseudo-reference images, paired with corresponding low-dose inputs, are fed into a robust regression network specifically designed for this task. To enhance feature extraction, we design a nested convolutional architecture using multi-scale features to improve the feature extraction capability of the regression Network. Furthermore, to strengthen the monotonic correlation between subjective and objective scores, we incorporate the relative distance information within each batch and enforce the relative ranking among the images. To propel research in this domain, we also build and release a new abdominal CT dataset with labeled IQA scores, tailored for benchmarking IQA methods. Extensive qualitative and quantitative experiments, conducted on both public datasets and our newly released dataset, demonstrate that our proposed PFGM-IQA method significantly outperforms state-of-the-art techniques in CT Image Quality Assessment. Haifeng Zhao 0001, Tianxia Yang, Shaojie Zhang 0002, Yanping Fu |
IJCNN | 5 |
| 2025 | TopoPoint: Enhance Topology Reasoning via Endpoint Detection in Autonomous DrivingabstractTopology reasoning, which unifies perception and structured reasoning, plays a vital role in understanding intersections for autonomous driving. However, its performance heavily relies on the accuracy of lane detection, particularly at connected lane endpoints. Existing methods often suffer from lane endpoints deviation, leading to incorrect topology construction. To address this issue, we propose TopoPoint, a novel framework that explicitly detects lane endpoints and jointly reasons over endpoints and lanes for robust topology reasoning. During training, we independently initialize point and lane query, and proposed Point-Lane Merge Self-Attention to enhance global context sharing through incorporating geometric distances between points and lanes as an attention mask . We further design Point-Lane Graph Convolutional Network to enable mutual feature aggregation between point and lane query. During inference, we introduce Point-Lane Geometry Matching algorithm that computes distances between detected points and lanes to refine lane endpoints, effectively mitigating endpoint deviation. Extensive experiments on the OpenLane-V2 benchmark demonstrate that TopoPoint achieves state-of-the-art performance in topology reasoning (48.8 on OLS). Additionally, we propose DET$_p$ to evaluate endpoint detection, under which our method significantly outperforms existing approaches (52.6 v.s. 45.2 on DET$_p$). The code is released at https://github.com/Franpin/TopoPoint. Yanping Fu, Xinyuan Liu 0003, Tianyu Li 0004, Yike Ma |
NeurIPS | 1 |
| 2025 | Fully Automated SAM for Single-source Domain Generalization in Medical Image Segmentation
Huanli Zhuo, Leilei Ma 0002, Haifeng Zhao 0001, Dengdi Sun, Yanping Fu |
SMC | 6 |
| 2025 | Feature-Enhanced Multimodal Interaction model for emotion recognition in conversation
Yanping Fu, Xiaoyuan Yan, Jun Zhang 0096 |
Knowl. Based Syst. | 1 |
| 2025 | Single image shadow removal using 2D signed distance field
Yanping Fu, Dengdi Sun, Shaojie Zhang 0002, Haifeng Zhao 0001 |
Vis. Comput. | 1 |
| 2024 | Single Image Reflection removal Using Feature Difference EnhancementabstractMost existing reflection removal methods pay too much attention to the transmission layer and ignore the mutual complementary mechanisms between the transmission layer and the reflection layer. To make full use of the complementarity and distinction between the reflection and transmission layers in reflection-contaminated images, we propose a novel single image reflection removal framework using the feature difference and adaptive information exchange between the transmission and reflection layers of reflection-contaminated images. First, we design a feature difference enhancement module to distinguish and enhance the feature difference of the reflective and transmissive layers. Second, we propose an adaptive information exchange module between transmission and reflection layers in the decoder, which can capture more complementary information. Finally, we introduce a 1/4 selective Instance Normalization strategy to improve our reflection removal tasks further. The experimental results demonstrate the efficiency of the proposed method and superior performance against state-of-the-art methods. Haifeng Zhao 0001, Shaojie Zhang 0002, Yanping Fu |
ICASSP | 4 |
| 2024 | Text-Region Matching for Multi-Label Image Recognition with Missing LabelsabstractRecently, large-scale visual language pre-trained (VLP) models have demonstrated impressive performance across various downstream tasks. Motivated by these advancements, pioneering efforts have emerged in multi-label image recognition with missing labels, leveraging VLP prompt-tuning technology. However, they usually cannot match text and vision features well, due to complicated semantics gaps and missing labels in a multi-label image. To tackle this challenge, we propose Text-Region Matching for optimizing Multi-Label prompt tuning, namely TRM-ML, a novel method for enhancing meaningful cross-modal matching. Compared to existing methods, we advocate exploring the information of category-aware regions rather than the entire image or pixels, which contributes to bridging the semantic gap between textual and visual representations in a one-to-one matching manner. Concurrently, we further introduce multimodal contrastive learning to narrow the semantic gap between textual and visual modalities and establish intra-class and inter-class relationships. Additionally, to deal with missing labels, we propose a multimodal category prototype that leverages intra- and inter-category semantic relationships to estimate unknown labels, facilitating pseudo-label generation. Extensive experiments on the MS-COCO, PASCAL VOC, Visual Genome, NUS-WIDE, and CUB-200-211 benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art methods by a significant margin. Our code is available here. Leilei Ma 0002, Hongxing Xie, Lei Wang 0095, Yanping Fu, Dengdi Sun, Haifeng Zhao 0001 |
ACM Multimedia | 4 |
| 2024 | TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving ScenesabstractAs an emerging task that integrates perception and reasoning, topology reasoning in autonomous driving scenes has recently garnered widespread attention. However, existing work often emphasizes "perception over reasoning": they typically boost reasoning performance by enhancing the perception of lanes and directly adopt vanilla MLPs to learn lane topology from lane query. This paradigm overlooks the geometric features intrinsic to the lanes themselves and are prone to being influenced by inherent endpoint shifts in lane detection.
To tackle this issue, we propose an interpretable method for lane topology reasoning based on lane geometric distance and lane query similarity, named TopoLogic. This method mitigates the impact of endpoint shifts in geometric space, and introduces explicit similarity calculation in semantic space as a complement. By integrating results from both spaces, our methods provides more comprehensive information for lane topology. Ultimately, our approach significantly outperforms the existing state-of-the-art methods on the mainstream benchmark OpenLane-V2 (23.9 v.s. 10.9 in TOP$_{ll}$ and 44.1 v.s. 39.8 in OLS on subsetA). Additionally, our proposed geometric distance topology reasoning method can be incorporated into well-trained models without re-training, significantly enhancing the performance of lane topology reasoning. The code is released at https://github.com/Franpin/TopoLogic. Yanping Fu, Wenbin Liao, Xinyuan Liu 0003, Yike Ma |
NeurIPS | 1 |
| 2024 | Domain Adaptive Lung Nodule Detection in X-Ray ImageabstractMedical images from different healthcare centers exhibit varied data distributions, posing significant challenges for adapting lung nodule detection due to the domain shift between training and application phases. Traditional unsupervised domain adaptive detection methods often struggle with this shift, leading to suboptimal outcomes. To overcome these challenges, we introduce a novel domain adaptive approach for lung nodule detection that leverages mean teacher self-training and contrastive learning. First, we propose a hierarchical contrastive learning strategy to refine nodule representations and enhance the distinction between nodules and background. Second, we introduce a nodule-level domain-invariant feature learning (NDL) module to capture domain-invariant features through adversarial learning across different domains. Additionally, we propose a new annotated dataset of X-ray images to aid in advancing lung nodule detection research. Extensive experiments conducted on multiple X-ray datasets demonstrate the efficacy of our approach in mitigating domain shift impacts. Haifeng Zhao 0001, Lixiang Jiang, Leilei Ma 0002, Dengdi Sun, Yanping Fu |
SMC | 5 |
| 2024 | Hybrid cross-modal interaction learning for multimodal sentiment analysis
Yanping Fu, Ruidi Yang, Cuiyou Yao |
Neurocomputing | 1 |
| 2024 | DGECN++: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation via Attention MechanismabstractMonocular object 6D pose estimation is a fundamental yet challenging task in computer vision. Recently, deep learning has been proven to be capable of predicting remarkable results in this task. Existing works often adopt a two-stage pipeline with establishing 2D-3D correspondences and utilizing a PnP/RANSAC or differentiable PnP algorithm to recover 6 degrees-of-freedom (6DoF) pose parameters. However, most of them hardly consider the geometric features in 3D space, and ignore the topological cues when performing differentiable PnP algorithms. To this end, we present an improved end-to-end monocular 6D pose estimation method (DGECN++) that incorporates depth estimation and a geometric-aware learnable PnP network. Our method is based on keypoints. First we detect the 2D keypoints that correspond to the 3D model. We then integrate differentiable PnP/RANSAC algorithm to create an end-to-end pipeline for 6D pose estimation. We focuses on the following three key aspects: 1) We utilize the estimated depth information to guide the process of extracting 2D-3D correspondences and refine the results using a cascaded differentiable PnP/RANSAC algorithm that incorporates geometric information. 2) We leverage the uncertainty of the estimated depth map to enhance the accuracy and robustness of the predicted 6D pose. 3) We propose a differentiable Perspective-n-Point (PnP) algorithm based on edge convolution and self-attention to explore the topological relationships between 2D-3D correspondences. Experimental results demonstrate that our proposed network surpasses existing methods in terms of both effectiveness and efficiency. Tuo Cao, Yanping Fu, Shengjie Zheng, Fei Luo 0004, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Seamless Texture Optimization for RGB-D ReconstructionabstractRestoring high-fidelity textures for 3D reconstructed models are an increasing demand in AR/VR, cultural heritage protection, entertainment, and other relevant fields. Due to geometric errors and camera pose drifting, existing texture mapping algorithms are either plagued by blurring and ghosting or suffer from undesirable visual seams. In this paper, we propose a novel tri-directional similarity texture synthesis method to eliminate the texture inconsistency in RGB-D 3D reconstruction and generate visually realistic texture mapping results. In addition to RGB color information, we incorporate a novel color image texture detail layer serving as an additional context to improve the effectiveness and robustness of the proposed method. First, we select an optimal texture image for each triangle face of the reconstructed model to avoid texture blurring and ghosting. During the selection procedure, the texture details are weighted to avoid generating texture chart partitions across high-frequency areas. Then, we optimize the camera pose of each texture image to align with the reconstructed 3D shape. Next, we propose a tri-directional similarity function to resynthesize the image context within the boundary stripe of texture charts, which can significantly diminish the occurrence of texture seams. Finally, we introduce a global color harmonization method to address the color inconsistency between texture images captured from different viewpoints. The experimental results demonstrate that the proposed method outperforms state-of-the-art texture mapping methods and effectively overcomes texture tearing, blurring, and ghosting artifacts. Yanping Fu, Qingan Yan, Huajian Zhou, Jin Tang 0001, Chunxia Xiao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Sparse RGB-D images create a real thing: A flexible voxel based 3D reconstruction pipeline for single objectabstractReconstructing 3D models for single objects with complex backgrounds has wide applications like 3D printing, AR/VR, and so on. It is necessary to consider the tradeoff between capturing data at low cost and getting high-quality reconstruction results. In this work, we propose a voxel-based modeling pipeline with sparse RGB-D images to effectively and efficiently reconstruct a single real object without the geometrical post-processing operation on background removal. First, referring to the idea of VisualHull, useless and inconsistent voxels of a targeted object are clipped. It helps focus on the target object and rectify the voxel projection information. Second, a modified TSDF calculation and voxel filling operations are proposed to alleviate the problem of depth missing in the depth images. They can improve TSDF value completeness for voxels on the surface of the object. After the mesh is generated by the MarchingCube, texture mapping is optimized with view selection, color optimization, and camera parameters fine-tuning. Experiments on Kinect capturing dataset, TUM public dataset, and virtual environment dataset validate the effectiveness and flexibility of our proposed pipeline. Fei Luo 0004, Yongqiong Zhu, Yanping Fu, Huajian Zhou, Zezheng Chen, Chunxia Xiao |
Vis. Informatics | 3 |
| 2022 | DGECN: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose EstimationabstractMonocular 6D pose estimation is a fundamental task in computer vision. Existing works often adopt a two-stage pipeline by establishing correspondences and utilizing a RANSAC algorithm to calculate 6 degrees-of-freedom (6DoF) pose. Recent works try to integrate differentiable RANSAC algorithms to achieve an end-to-end 6D pose estimation. However, most of them hardly consider the geometric features in 3D space, and ignore the topology cues when performing differentiable RANSAC algorithms. To this end, we proposed a Depth-Guided Edge Convolutional Network (DGECN) for 6D pose estimation task. We have made efforts from the following three aspects: 1) We take advantages of estimated depth information to guide both the correspondences-extraction process and the cascaded differentiable RANSAC algorithm with geometric information. 2) We leverage the uncertainty of the estimated depth map to improve accuracy and robustness of the output 6D pose. 3) We propose a differentiable Perspective-n-Point(PnP) algorithm via edge convolution to explore the topology relations between 2D-3D correspondences. Experiments demonstrate that our proposed network outperforms current works on both effectiveness and efficiency. Tuo Cao, Fei Luo 0004, Yanping Fu, Shengjie Zheng, Chunxia Xiao |
CVPR | 3 |
| 2022 | Depth-Aware Shadow RemovalabstractAbstract Shadow removal from a single image is an ill‐posed problem because shadow generation is affected by the complex interactions of geometry, albedo, and illumination. Most recent deep learning‐based methods try to directly estimate the mapping between the non‐shadow and shadow image pairs to predict the shadow‐free image. However, they are not very effective for shadow images with complex shadows or messy backgrounds. In this paper, we propose a novel end‐to‐end depth‐aware shadow removal method without using depth images, which estimates depth information from RGB images and leverages the depth feature as guidance to enhance shadow removal and refinement. The proposed framework consists of three components, including depth prediction, shadow removal, and boundary refinement. First, the depth prediction module is used to predict the corresponding depth map of the input shadow image. Then, we propose a new generative adversarial network (GAN) method integrated with depth information to remove shadows in the RGB image. Finally, we propose an effective boundary refinement framework to alleviate the artifact around boundaries after shadow removal by depth cues. We conduct experiments on several public datasets and real‐world shadow images. The experimental results demonstrate the efficiency of the proposed method and superior performance against state‐of‐the‐art methods. Yanping Fu, Zhenyu Gai, Haifeng Zhao 0001, Shaojie Zhang 0002, Ying Shan, Yang Wu 0001, Jin Tang 0001 |
Comput. Graph. Forum | 1 |
| 2022 | Domain adaptation with a shrinkable discrepancy strategy for cross-domain sentiment classification
Yanping Fu |
Neurocomputing | 1 |
| 2022 | Contrastive transformer based domain adaptation for multi-source cross-domain sentiment classification
Yanping Fu |
Knowl. Based Syst. | 1 |
| 2022 | NeuralRoom: Geometry-Constrained Neural Implicit Surfaces for Indoor Scene ReconstructionabstractWe present a novel neural surface reconstruction method called NeuralRoom for reconstructing room-sized indoor scenes directly from a set of 2D images. Recently, implicit neural representations have become a promising way to reconstruct surfaces from multiview images due to their high-quality results and simplicity. However, implicit neural representations usually cannot reconstruct indoor scenes well because they suffer severe shape-radiance ambiguity. We assume that the indoor scene consists of texture-rich and flat texture-less regions. In texture-rich regions, the multiview stereo can obtain accurate results. In the flat area, normal estimation networks usually obtain a good normal estimation. Based on the above observations, we reduce the possible spatial variation range of implicit neural surfaces by reliable geometric priors to alleviate shape-radiance ambiguity. Specifically, we use multiview stereo results to limit the NeuralRoom optimization space and then use reliable geometric priors to guide NeuralRoom training. Then the NeuralRoom would produce a neural scene representation that can render an image consistent with the input training images. In addition, we propose a smoothing method called perturbation-residual restrictions to improve the accuracy and completeness of the flat region, which assumes that the sampling points in a local surface should have the same normal and similar distance to the observation center. Experiments on the ScanNet dataset show that our method can reconstruct the texture-less area of indoor scenes while maintaining the accuracy of detail. We also apply NeuralRoom to more advanced multiview reconstruction algorithms and significantly improve their reconstruction quality. Yusen Wang 0002, Zongcheng Li, Yu Jiang 0007, Kaixuan Zhou, Tuo Cao, Yanping Fu, Chunxia Xiao |
ACM Trans. Graph. | 6 |
| 2022 | CGFNet: cross-guided fusion network for RGB-thermal semantic segmentation
Yanping Fu, Qiaoqiao Chen, Haifeng Zhao 0001 |
Vis. Comput. | 1 |
| 2021 | Adaptive depth estimation for pyramid multi-view stereo
Yanping Fu, Qingan Yan, Fei Luo 0004, Chunxia Xiao |
Comput. Graph. | 2 |
| 2021 | CGSPN : cascading gated self-attention and phrase-attention network for sentence modeling
Yanping Fu, Yun Liu 0001 |
J. Intell. Inf. Syst. | 1 |
| 2021 | Dense multiview stereo based on image texture enhancementabstractAbstract In this paper, we propose a novel Multiview Stereo (MVS) method which can effectively estimate geometry in low‐textured regions. Conventional MVS algorithms predict geometry by performing dense correspondence estimation across multiple views under the constraint of epipolar geometry. As low‐textured regions contain less feature information for reliable matching, estimating geometry for low‐textured regions remains hard work for previous MVS methods. To address this issue, we propose an MVS method based on texture enhancement. By enhancing texture information for each input image via our multiscale bilateral decomposition and reconstruction algorithm, our method can estimate reliable geometry for low‐textured regions that are intractable for previous MVS methods. To densify the final output point cloud, we further propose a novel selective joint bilateral propagation filter, which can effectively propagate reliable geometry estimation to neighboring unpredicted regions. We validate the effectiveness of our method on the ETH3D benchmark. Quantitative and qualitative comparisons demonstrate that our method can significantly improve the quality of reconstruction in low‐textured regions. Mengqiang Wei, Yanping Fu, Qingan Yan, Chunxia Xiao |
Comput. Animat. Virtual Worlds | 3 |
| 2021 | Cross-domain sentiment classification based on key pivot and non-pivot extraction
Yanping Fu, Yun Liu 0001 |
Knowl. Based Syst. | 1 |
| 2020 | Joint Texture and Geometry Optimization for RGB-D ReconstructionabstractDue to inevitable noises and quantization error, the reconstructed 3D models via RGB-D sensors always accompany geometric error and camera drifting, which consequently lead to blurring and unnatural texture mapping results. Most of the 3D reconstruction methods focus on either geometry refinement or texture improvement respectively, which subjectively decouples the inter-relationship between geometry and texture. In this paper, we propose a novel approach that can jointly optimize the camera poses, texture and geometry of the reconstructed model, and color consistency between the key-frames. Instead of computing Shape-From-Shading (SFS) expensively, our method directly optimizes the reconstructed mesh according to color and geometric consistency and high-boost normal cues, which can effectively overcome the texture-copy problem generated by SFS and achieve more detailed shape reconstruction. As the joint optimization involves multiple correlated terms, therefore, we further introduce an iterative framework to interleave the optimal state. The experiments demonstrate that our method can recover not only fine-scale geometry but also high-fidelity texture. Yanping Fu, Qingan Yan, Chunxia Xiao |
CVPR | 1 |
| 2020 | Folding patch correspondence for multiview stereoabstractAbstract In this article, we propose the novel folding patch model which can replace the traditional patch model utilized in patch‐based multiview stereo (MVS) methods to significantly improve the reconstruction results. The patch model is applied as an approximation of the scene surface differential in the geometric estimation procedure. By minimizing the photometric discrepancy of the projection of the patch model on multiple source images, patch‐based MVS algorithms optimize the position and normal values for the 3D hypothesis of the target pixel. The optimization is based on the assumption that the patch model can fit the target scene surface perfectly. However, when it comes to complex scenes crowded with sharp edges, splintery surfaces, or round surfaces, the patch model is inherently not suitable since even from the microscopic perspective these surfaces are not entirely flat. We construct the folding patch model by folding the traditional patch model from the middle line. By adjusting the folding angle and direction, the folding patch model can fit complex surfaces more flexibly. We apply our folding patch model to the representative open‐source patch based multiview stereo (PMVS) and COLMAP, and validate the effectiveness on ETH3D benchmark and data sets captured in nature. The results demonstrate that utilizing the folding patch model can significantly improve the behavior of PMVS and COLMAP, especially on data sets mainly consist of complex surfaces from plants. Yanping Fu, Qingan Yan, Chunxia Xiao |
Comput. Animat. Virtual Worlds | 2 |
| 2020 | Transparent object segmentation from casually captured videosabstractAbstract Segmentation of transparent objects from sequences can be very useful in computer vision applications. However, without additional auxiliary information it can be hard work for traditional segmentation methods, as light in the transparent area captured by RGB cameras mostly derive from the background and the appearance of transparent objects changes with surroundings. In this article, we present a from‐coarse‐to‐fine transparent object segmentation method, which utilizes trajectory clustering to roughly distinguish the transparent from the background and refine the segmentation based on combination information of color and distortion. We further incorporate the transparency saliency with color and trajectory smoothness throughout the video to acquire a spatiotemporal segmentation based on graph‐cut. We conduct our method on various datasets. The results demonstrate that our method can successfully segment transparent objects from the background. Yanping Fu, Qingan Yan, Chunxia Xiao |
Comput. Animat. Virtual Worlds | 2 |
| 2020 | Real-time dense 3D reconstruction and camera tracking via embedded planes representation
Yanping Fu, Qingan Yan, Alix L. H. Chow, Chunxia Xiao |
Vis. Comput. | 1 |
| 2019 | Community Detection with Indirect Neighbors based on Granular Computing in Social NetworksabstractSocial relations exist in real life widely, which can be abstracted into graphs to show various social networks. Community detection in social networks has become an important and effective methodology to understand the structure and function of real world networks. In this paper, the algorithm we proposed fully considers the influence of direct nodes and indirect nodes and expresses the relationship between users comprehensively. The algorithm runs on the social network graph based on a new granular computing method. Experimental results on bench mark data show the superiority of the proposed algorithm compared to other well-known methods. Naiyue Chen, Yanping Fu |
IJCNN | 6 |
| 2019 | Pyramid Multi-View Stereo with Local ConsistencyabstractAbstract In this paper, we propose a PatchMatch‐based Multi‐View Stereo (MVS) algorithm which can efficiently estimate geometry for the textureless area. Conventional PatchMatch‐based MVS algorithms estimate depth and normal hypotheses mainly by optimizing photometric consistency metrics between patch in the reference image and its projection on other images. The photometric consistency works well in textured regions but can not discriminate textureless regions, which makes geometry estimation for textureless regions hard work. To address this issue, we introduce the local consistency. Based on the assumption that neighboring pixels with similar colors likely belong to the same surface and share approximate depth‐normal values, local consistency guides the depth and normal estimation with geometry from neighboring pixels with similar colors. To fasten the convergence of pixelwise local consistency across the image, we further introduce a pyramid architecture similar to previous work which can also provide coarse estimation at upper levels. We validate the effectiveness of our method on the ETH3D benchmark and Tanks and Temples benchmark. Results show that our method outperforms the state‐of‐the‐art. Yanping Fu, Qingan Yan, Chunxia Xiao |
Comput. Graph. Forum | 2 |
| 2018 | Texture Mapping for 3D Reconstruction With RGB-D SensorabstractAcquiring realistic texture details for 3D models is important in 3D reconstruction. However, the existence of geometric errors, caused by noisy RGB-D sensor data, always makes the color images cannot be accurately aligned onto reconstructed 3D models. In this paper, we propose a global-to-local correction strategy to obtain more desired texture mapping results. Our algorithm first adaptively selects an optimal image for each face of the 3D model, which can effectively remove blurring and ghost artifacts produced by multiple image blending. We then adopt a non-rigid global-to-local correction step to reduce the seaming effect between textures. This can effectively compensate for the texture and the geometric misalignment caused by camera pose drift and geometric errors. We evaluate the proposed algorithm in a range of complex scenes and demonstrate its effective performance in generating seamless high fidelity textures for 3D models. Yanping Fu, Qingan Yan, Long Yang 0001, Chunxia Xiao |
CVPR | 1 |
| 2018 | Surface Reconstruction via Fusing Sparse-Sequence of Depth ImagesabstractHandheld scanning using commodity depth cameras provides a flexible and low-cost manner to get 3D models. The existing methods scan a target by densely fusing all the captured depth images, yet most frames are redundant. The jittering frames inevitably embedded in handheld scanning process will cause feature blurring on the reconstructed model and even trigger the scan failure (i.e., camera tracking losing). To address these problems, in this paper, we propose a novel sparse-sequence fusion (SSF) algorithm for handheld scanning using commodity depth cameras. It first extracts related measurements for analyzing camera motion. Then based on these measurements, we progressively construct a supporting subset for the captured depth image sequence to decrease the data redundancy and the interference from jittering frames. Since SSF will reveal the intrinsic heavy noise of the original depth images, our method introduces a refinement process to eliminate the raw noise and recover geometric features for the depth images selected into the supporting subset. We finally obtain the fused result by integrating the refined depth images into the truncated signed distance field (TSDF) of the target. Multiple comparison experiments are conducted and the results verify the feasibility and validity of SSF for handheld scanning with a commodity depth camera. Long Yang 0001, Qingan Yan, Yanping Fu, Chunxia Xiao |
IEEE Trans. Vis. Comput. Graph. | 3 |