EDBT 2026 Demo / reviewers in the wild / expert
Youcheng Cai
dblp:278/2135
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-0599-8418ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MLENet: Multi-level efficient network based on single-scale feature extraction for human keypoint estimation
Dong Wang 0043, Youcheng Cai, Yiming Tang 0001, Wenjun Xie, Xiaoping Liu 0003 |
Expert Syst. Appl. | 2 |
| 2026 | AHC-NeRF: Autonomous, High-Quality Neural Reconstruction of Two-Layer Complex Nested Transparent ObjectsabstractReconstructing transparent objects with high fidelity presents significant challenges due to complex light refraction and reflection. Existing methods rely on intentionally designed patterns observed behind the transparent object to infer the correspondence between rays and the background, thereby improving the precision of the reconstruction. However, they are hindered by a refraction-tracing-based strategy that fails to reconstruct complex nested transparent objects and a tedious view-capture strategy relying on images captured from empirically determined viewpoints. To overcome these obstacles, we propose AHC-NeRF, an autonomous, high-quality neural SDF-based framework designed for reconstructing two-layer complex nested transparent objects. Firstly, our framework combines neural SDF with single-pixel imaging, a reflection-based method, which utilizes point-pair priors as guidance to achieve high-quality reconstruction of both the outer and inner surfaces. Secondly, we propose an adaptive single-pixel imaging method that achieves an acceleration of 1-2 orders of magnitude compared to vanilla single-pixel imaging for the acquisition of point-pair priors. Finally, we introduce a novel view-planning strategy that progressively identifies the viewpoints with the highest information gain throughout the optimization process, thereby achieving high-quality surface reconstruction. Extensive experimental results on both synthetic and real-world datasets demonstrate that AHC-NeRF outperforms state-of-the-art methods. Youcheng Cai, Li Li 0094, Ligang Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | Hash-NURF: efficient nested transparent object reconstruction using multi-resolution hash encoding
Youcheng Cai |
Vis. Comput. | 3 |
| 2025 | Importance Sampling Guided Neural Radiosity
Huangsheng Du, Youcheng Cai, Yutian Zhu |
Comput. Graph. | 2 |
| 2025 | Computational multi-layered wood carving art
Zhi Li 0076, Youcheng Cai, Xiaoya Zhai, Ketian Zhang, Ligang Liu 0001, Yi Min Xie, Xiao-Ming Fu 0001 |
Comput. Graph. | 4 |
| 2025 | DragMV: Interactive drag-style manipulation on multi-view images
Junyan Qi, Youcheng Cai |
Comput. Graph. | 2 |
| 2025 | Uanet: uncertainty-aware cost volume aggregation-based multi-view stereo for 3D reconstruction
Youcheng Cai, Jiale Yang |
Vis. Comput. | 2 |
| 2024 | MV2MV: Multi-View Image Translation via View-Consistent Diffusion ModelsabstractImage translation has various applications in computer graphics and computer vision, aiming to transfer images from one domain to another. Thanks to the excellent generation capability of diffusion models, recent single-view image translation methods achieve realistic results. However, directly applying diffusion models for multi-view image translation remains challenging for two major obstacles: the need for paired training data and the limited view consistency. To overcome the obstacles, we present a first unified multi-view image to multi-view image translation framework based on diffusion models, called MV2MV. Firstly, we propose a novel self-supervised training strategy that exploits the success of off-the-shelf single-view image translators and the 3D Gaussian Splatting (3DGS) technique to generate pseudo ground truths as supervisory signals, leading to enhanced consistency and fine details. Additionally, we propose a latent multi-view consistency block, which utilizes the latent-3DGS as the underlying 3D representation to facilitate information exchange across multi-view images and inject 3D prior into the diffusion model to enforce consistency. Finally, our approach simultaneously optimizes the diffusion model and 3DGS to achieve a better trade-off between consistency and realism. Extensive experiments across various translation tasks demonstrate that MV2MV outperforms task-specific specialists in both quantitative and qualitative. Youcheng Cai, Runshi Li, Ligang Liu 0001 |
ACM Trans. Graph. | 1 |
| 2023 | MFNet: Multi-level fusion aware feature pyramid based multi-view stereo network for 3D reconstruction
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xiaoping Liu 0003 |
Appl. Intell. | 1 |
| 2023 | Transformer-based rapid human pose estimation network
Dong Wang 0043, Wenjun Xie, Youcheng Cai, Xinjie Li 0006, Xiaoping Liu 0003 |
Comput. Graph. | 3 |
| 2023 | HTMatch: An efficient hybrid transformer based graph neural network for local feature matching
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xinjie Li 0006, Xiaoping Liu 0003 |
Signal Process. | 1 |
| 2023 | GlcMatch: global and local constraints for reliable feature matching
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xintao Huang, Xiaoping Liu 0003 |
Vis. Comput. | 1 |
| 2022 | A Fast and Effective Transformer for Human Pose EstimationabstractMost of the existing human pose estimation methods improve accuracy by constantly increasing computational resources. However, balancing the efficiency and efficacy of the model is the key to enhancing the real application value. In this work, we present a Fast and Effective Transformer model to ensure the efficiency and efficacy of the model, called FET. Specifically, the FET consists of three parts: Feature Extraction Module (FEM), Feature Interaction Module (FIM) and Feature Decode Module (FDM). The FEM is used to efficiently extract low-level features from input images. Unlike CNN-based strategies, the FIM enables our model to capture global dependencies by self-attention, thus improving the accuracy for human pose estimation. The FDM is a multistage way that gradually recovers the size of the features to obtain a higher-quality target heatmap. In addition, Feature Squeeze Attention is introduced in the FET to further improve the overall performance of our model. Extensive experiments show that our method is 1.7× and 7× faster than SimpleBaseline and HRNet-32, respectively, while achieving comparable or even better results with the most state-of-the-art methods on the COCO dataset and the MPII dataset. Dong Wang 0043, Wenjun Xie, Youcheng Cai, Xiaoping Liu 0003 |
IEEE Signal Process. Lett. | 3 |