Youcheng Cai

dblp:278/2135 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-0599-8418ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MLENet: Multi-level efficient network based on single-scale feature extraction for human keypoint estimation
Dong Wang 0043, Youcheng Cai, Yiming Tang 0001, Wenjun Xie, Xiaoping Liu 0003
Expert Syst. Appl.2
2026 AHC-NeRF: Autonomous, High-Quality Neural Reconstruction of Two-Layer Complex Nested Transparent Objects
abstract
Reconstructing transparent objects with high fidelity presents significant challenges due to complex light refraction and reflection. Existing methods rely on intentionally designed patterns observed behind the transparent object to infer the correspondence between rays and the background, thereby improving the precision of the reconstruction. However, they are hindered by a refraction-tracing-based strategy that fails to reconstruct complex nested transparent objects and a tedious view-capture strategy relying on images captured from empirically determined viewpoints. To overcome these obstacles, we propose AHC-NeRF, an autonomous, high-quality neural SDF-based framework designed for reconstructing two-layer complex nested transparent objects. Firstly, our framework combines neural SDF with single-pixel imaging, a reflection-based method, which utilizes point-pair priors as guidance to achieve high-quality reconstruction of both the outer and inner surfaces. Secondly, we propose an adaptive single-pixel imaging method that achieves an acceleration of 1-2 orders of magnitude compared to vanilla single-pixel imaging for the acquisition of point-pair priors. Finally, we introduce a novel view-planning strategy that progressively identifies the viewpoints with the highest information gain throughout the optimization process, thereby achieving high-quality surface reconstruction. Extensive experimental results on both synthetic and real-world datasets demonstrate that AHC-NeRF outperforms state-of-the-art methods.
Youcheng Cai, Li Li 0094, Ligang Liu 0001
IEEE Trans. Vis. Comput. Graph.1
2026 Hash-NURF: efficient nested transparent object reconstruction using multi-resolution hash encoding
Youcheng Cai
Vis. Comput.3
2025 Importance Sampling Guided Neural Radiosity
Huangsheng Du, Youcheng Cai, Yutian Zhu
Comput. Graph.2
2025 Computational multi-layered wood carving art
Zhi Li 0076, Youcheng Cai, Xiaoya Zhai, Ketian Zhang, Ligang Liu 0001, Yi Min Xie, Xiao-Ming Fu 0001
Comput. Graph.4
2025 DragMV: Interactive drag-style manipulation on multi-view images
Junyan Qi, Youcheng Cai
Comput. Graph.2
2025 Uanet: uncertainty-aware cost volume aggregation-based multi-view stereo for 3D reconstruction
Youcheng Cai, Jiale Yang
Vis. Comput.2
2024 MV2MV: Multi-View Image Translation via View-Consistent Diffusion Models
abstract
Image translation has various applications in computer graphics and computer vision, aiming to transfer images from one domain to another. Thanks to the excellent generation capability of diffusion models, recent single-view image translation methods achieve realistic results. However, directly applying diffusion models for multi-view image translation remains challenging for two major obstacles: the need for paired training data and the limited view consistency. To overcome the obstacles, we present a first unified multi-view image to multi-view image translation framework based on diffusion models, called MV2MV. Firstly, we propose a novel self-supervised training strategy that exploits the success of off-the-shelf single-view image translators and the 3D Gaussian Splatting (3DGS) technique to generate pseudo ground truths as supervisory signals, leading to enhanced consistency and fine details. Additionally, we propose a latent multi-view consistency block, which utilizes the latent-3DGS as the underlying 3D representation to facilitate information exchange across multi-view images and inject 3D prior into the diffusion model to enforce consistency. Finally, our approach simultaneously optimizes the diffusion model and 3DGS to achieve a better trade-off between consistency and realism. Extensive experiments across various translation tasks demonstrate that MV2MV outperforms task-specific specialists in both quantitative and qualitative.
Youcheng Cai, Runshi Li, Ligang Liu 0001
ACM Trans. Graph.1
2023 MFNet: Multi-level fusion aware feature pyramid based multi-view stereo network for 3D reconstruction
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xiaoping Liu 0003
Appl. Intell.1
2023 Transformer-based rapid human pose estimation network
Dong Wang 0043, Wenjun Xie, Youcheng Cai, Xinjie Li 0006, Xiaoping Liu 0003
Comput. Graph.3
2023 HTMatch: An efficient hybrid transformer based graph neural network for local feature matching
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xinjie Li 0006, Xiaoping Liu 0003
Signal Process.1
2023 GlcMatch: global and local constraints for reliable feature matching
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xintao Huang, Xiaoping Liu 0003
Vis. Comput.1
2022 A Fast and Effective Transformer for Human Pose Estimation
abstract
Most of the existing human pose estimation methods improve accuracy by constantly increasing computational resources. However, balancing the efficiency and efficacy of the model is the key to enhancing the real application value. In this work, we present a Fast and Effective Transformer model to ensure the efficiency and efficacy of the model, called FET. Specifically, the FET consists of three parts: Feature Extraction Module (FEM), Feature Interaction Module (FIM) and Feature Decode Module (FDM). The FEM is used to efficiently extract low-level features from input images. Unlike CNN-based strategies, the FIM enables our model to capture global dependencies by self-attention, thus improving the accuracy for human pose estimation. The FDM is a multistage way that gradually recovers the size of the features to obtain a higher-quality target heatmap. In addition, Feature Squeeze Attention is introduced in the FET to further improve the overall performance of our model. Extensive experiments show that our method is 1.7× and 7× faster than SimpleBaseline and HRNet-32, respectively, while achieving comparable or even better results with the most state-of-the-art methods on the COCO dataset and the MPII dataset.
Dong Wang 0043, Wenjun Xie, Youcheng Cai, Xiaoping Liu 0003
IEEE Signal Process. Lett.3