EDBT 2026 Demo / reviewers in the wild / expert
Yifan Wang 0026
dblp:47/6959-26
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0008-7565-7839ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
1.7 | 2 | 2025 | Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation · ACM Trans. Graph. 2025 FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction · CVPR 2025 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.9 | 1 | 2025 | Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation · ACM Trans. Graph. 2025 |
Computer vision › 3D vision › 3d scene modeling › scene representation
3d scene representation |
0.9 | 1 | 2025 | FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction · CVPR 2025 |
Computer vision › 3D vision › novel view synthesis
dynamic view synthesis |
0.9 | 1 | 2025 | FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction · CVPR 2025 |
Computer vision › 3D vision
camera pose estimation |
0.8 | 1 | 2024 | Detector-Free Structure from Motion · CVPR 2024 |
Computer vision › 3D vision › feature matching › dense feature matching
detector-free matching |
0.8 | 1 | 2024 | Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed · CVPR 2024 |
Computer vision › 3D vision › feature matching › local feature matching
efficient feature matching |
0.8 | 1 | 2024 | Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed · CVPR 2024 |
Computer vision › 3D vision › feature matching
local feature matching |
0.8 | 1 | 2024 | Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed · CVPR 2024 |
Computer vision › 3D vision › pose estimation
multi-view pose estimation |
0.8 | 1 | 2024 | Detector-Free Structure from Motion · CVPR 2024 |
Computer vision › 3D vision › feature matching
semi-dense feature matching |
0.8 | 1 | 2024 | Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed · CVPR 2024 |
Computer vision › 3D vision
structure from motion |
0.8 | 1 | 2024 | Detector-Free Structure from Motion · CVPR 2024 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 1 | 2024 | Detector-Free Structure from Motion · CVPR 2024 |
Computer vision › 3D vision › 3d reconstruction
point cloud reconstruction |
0.2 | 1 | 2024 | Detector-Free Structure from Motion · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
streaming feature learning · 0.9gaussian splatting · 0.9differentiable rendering · 0.9deformation field · 0.9contrastive loss · 0.9two-stage correlation · 0.8iterative refinement · 0.8attention-based multi-view matching · 0.8aggregated attention · 0.8adaptive token selection · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Depth Foundation Models: Recent Trends in Vision-Based Depth EstimationabstractDepth estimation is a fundamental task in 3D computer vision, crucial for applications such as 3D reconstruction, free-viewpoint rendering, robotics, autonomous driving, and AR/VR technologies. Traditional methods relying on hardware sensors like LiDAR are often limited by their high costs, low resolution, and sensitivity to the environment, limiting their applicability to real-world scenarios. Recent advances in vision-based methods offer a promising alternative, yet they face challenges in generalization and stability due to either the low capacity of model architectures or reliance on domain-specific and small-scale datasets. The emergence of scaling laws and foundation models in other domains has inspired the development of “depth foundation models”: deep neural networks trained on large datasets with strong zeroshot generalization capabilities. This paper surveys the evolution of deep learning architectures and paradigms for depth estimation across monocular, stereo, multiview, and monocular video settings. We explore the potential of these models to address existing challenges and we also provide a comprehensive overview of large-scale datasets that can facilitate their development. By identifying key architectures and training strategies, we aim to highlight the path towards robust depth foundation models, offering insights for future research and applications. Zhen Xu 0008, Sida Peng, Haotong Lin, Jiahao Shao, Peishan Yang, Qinglin Yang, Sheng Miao, Yifan Wang 0026, Ruizhen Hu, Yiyi Liao, Xiaowei Zhou 0001, Hujun Bao |
Comput. Vis. Media | 11 |
| 2025 | FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene ReconstructionabstractThis paper addresses the challenge of reconstructing dynamic 3D scenes with complex motions. Some recent works define 3D Gaussian primitives in the canonical space and use deformation fields to map canonical primitives to observation spaces, achieving real-time dynamic view synthesis. However, these methods often struggle to handle scenes with complex motions due to the difficulty of optimizing deformation fields. To overcome this problem, we propose FreeTimeGS, a novel 4D representation that allows Gaussian primitives to appear at arbitrary time and locations. In contrast to canonical Gaussian primitives, our representation possesses the strong flexibility, thus improving the ability to model dynamic 3D scenes. In addition, we endow each Gaussian primitive with an motion function, allowing it to move to neighboring regions over time, which reduces the temporal redundancy. Experiments results on several datasets show that the rendering quality of our method outperforms recent methods by a large margin. The code will be released for reproducibility. Yifan Wang 0026, Peishan Yang, Zhen Xu 0008, Jiaming Sun 0002, Zhanhua Zhang, Hujun Bao, Sida Peng, Xiaowei Zhou 0001 |
CVPR | 1 |
| 2025 | Split4D: Decomposed 4D Scene Reconstruction Without Video SegmentationabstractThis paper addresses the problem of decomposed 4D scene reconstruction from multi-view videos. Recent methods achieve this by lifting video segmentation results to a 4D representation through differentiable rendering techniques. Therefore, they heavily rely on the quality of video segmentation maps, which are often unstable, leading to unreliable reconstruction results. To overcome this challenge, our key idea is to represent the decomposed 4D scene with the Freetime FeatureGS and design a streaming feature learning strategy to accurately recover it from per-image segmentation maps, eliminating the need for video segmentation. Freetime FeatureGS models the dynamic scene as a set of Gaussian primitives with learnable features and linear motion ability, allowing them to move to neighboring regions over time. We apply a contrastive loss to Freetime FeatureGS, forcing primitive features to be close or far apart based on whether their projections belong to the same instance in the 2D segmentation map. As our Gaussian primitives can move across time, it naturally extends the feature learning to the temporal dimension, achieving 4D segmentation. Furthermore, we sample observations for training in a temporally ordered manner, enabling the streaming propagation of features over time and effectively avoiding local minima during the optimization process. Experimental results on several datasets show that the reconstruction quality of our method outperforms recent methods by a large margin. Yongzhen Hu, Yihui Yang, Haotong Lin, Yifan Wang 0026, Junting Dong, Yifu Deng, Hujun Bao, Xiaowei Zhou 0001, Sida Peng |
ACM Trans. Graph. | 4 |
| 2024 | Detector-Free Structure from MotionabstractWe propose a structure-from-motion framework to recover accurate camera poses and point clouds from unordered images. Traditional SfM systems typically rely on the successful detection of repeatable keypoints across multiple views as the first step, which is difficult for texture-poor scenes, and poor keypoint detection may break down the whole SfM system. We propose a detector-free SfM framework to draw benefits from the recent success of detector-free matchers to avoid the early determination of keypoints, while solving the multi-view inconsistency issue of detector-free matchers. Specifically, our framework first reconstructs a coarse SfM model from quantized detector-free matches. Then, it refines the model by a novel iterative refinement pipeline, which iterates between an attention-based multi-view matching module to refine feature tracks and a geometry refinement module to improve the reconstruction accuracy. Experiments demonstrate that the proposed framework outperforms existing detector-based SfM systems on common benchmark datasets. We also collect a texture-poor SfM dataset to demonstrate the capa-bility of our framework to reconstruct texture-poor scenes. Based on this framework, we take the first place in Image Matching Challenge 2023 [9]. Project page: https://zju3dv.github.io/DetectorFreeSfM/. Jiaming Sun 0002, Yifan Wang 0026, Sida Peng, Qixing Huang, Hujun Bao, Xiaowei Zhou 0001 |
CVPR | 3 |
| 2024 | Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like SpeedabstractWe present a novel method for efficiently producing semi-dense matches across images. Previous detector-free matcher LoFTR has shown remarkable matching capability in handling large-viewpoint change and texture-poor scenarios but suffers from low efficiency. We revisit its design choices and derive multiple improvements for both efficiency and accuracy. One key observation is that performing the transformer over the entire feature map is redundant due to shared local information, therefore we propose an aggregated attention mechanism with adaptive token selection for efficiency. Furthermore, we find spatial variance exists in LoFTR's fine correlation module, which is adverse to matching accuracy. A novel two-stage correlation layer is proposed to achieve accurate subpixel correspon-dences for accuracy improvement. Our efficiency optimized model is ~ 2.5 x faster than LoFTR which can even surpass state-of-the-art efficient sparse matching pipeline Super-Point + LightGlue. Moreover, extensive experiments show that our method can achieve higher accuracy compared with competitive semi-dense matchers, with considerable efficiency benefits. This opens up exciting prospects for large-scale or latency-sensitive applications such as image retrieval and 3D reconstruction. Project page: https://zju3dv.github.io/efficientloftr/. Yifan Wang 0026, Sida Peng, Dongli Tan, Xiaowei Zhou 0001 |
CVPR | 1 |