VLDB 2026 Research / reviewers in the wild / expert
Haowen Sun 0004
dblp:220/8228-4
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-0466-191XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PointVDP: Learning view-dependent projection by fireworks rays for 3D point cloud segmentation
Yueqi Duan, Haowen Sun 0004, Ziwei Wang 0010, Jiwen Lu, Yap-Peng Tan |
Pattern Recognit. | 3 |
| 2026 | ReconX: Reconstruct Any Scene From Sparse Views With Video Diffusion ModelabstractAdvancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from sparse views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction problem as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction. Nevertheless, it is challenging to preserve 3D view consistency when directly generating video frames from pre-trained models. To address this issue, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of ReconX over state-of-the-art methods in terms of quality and generalizability. Fangfu Liu, Wenqiang Sun, Hanyang Wang 0003, Yikai Wang 0001, Haowen Sun 0004, Junliang Ye, Jun Zhang 0004, Yueqi Duan |
IEEE Trans. Image Process. | 5 |
| 2026 | Ambiguity-Aware Point Cloud Segmentation by Adaptive Margin Contrastive LearningabstractThis paper proposes an adaptive margin contrastive learning method for 3D semantic segmentation on point clouds. Most existing methods use equally penalized objectives, which ignore the per-point ambiguities and less discriminated features stemming from transition regions. However, as highly ambiguous points may be indistinguishable even for humans, their manually annotated labels are less reliable, and hard constraints over these points would lead to sub-optimal models. To address this, we first design AMContrast3D, a method comprising contrastive learning into an ambiguity estimation framework, tailored to adaptive objectives for individual points based on ambiguity levels. As a result, our method promotes model training, which ensures the correctness of low-ambiguity points while allowing mistakes for high-ambiguity points. As ambiguities are formulated based on position discrepancies across labels, optimization during inference is constrained by the assumption that all unlabeled points are uniformly unambiguous, lacking ambiguity awareness. Inspired by the insight of joint training, we further propose AMContrast3D++ integrating with two branches trained in parallel, where a novel ambiguity prediction module concurrently learns point ambiguities from generated embeddings. To this end, we design a masked refinement mechanism that leverages predicted ambiguities to enable the ambiguous embeddings to be more reliable, thereby boosting segmentation performance and enhancing robustness. Experimental results on 3D indoor scene datasets, S3DIS and ScanNet, demonstrate the effectiveness of the proposed method. Code is available athttps://github.com/YangChenApril/AMContrast3D. Yueqi Duan, Haowen Sun 0004, Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Multim. | 3 |
| 2024 | MirageRoom: 3D Scene Segmentation with 2D Pre-Trained Models by Mirage ProjectionabstractNowadays, leveraging 2D images and pre-trained mod- els to guide 3D point cloud feature representation has shown a remarkable potential to boost the performance of 3D fundamental models. While some works rely on additional data such as 2D real-world images and their corre- sponding camera poses, recent studies target at using point cloud exclusively by designing 3D-to-2D projection. How- ever, in the indoor scene scenario, existing 3D-to-2D pro- jection strategies suffer from severe occlusions and incoher- ence, which fail to contain sufficient information for fine- grained point cloud segmentation task. In this paper, we ar- gue that the crux of the matter resides in the basic premise of existing projection strategies that the medium is homo- geneous, thereby projection rays propagate along straight lines and behind objects are occluded by front ones. In- spired by the phenomenon of mirage where the occluded objects are exposed by distorted light rays due to heteroge- neous medium refraction rate, we propose MirageRoom by designing parametric mirage projection with heterogeneous medium to obtain series of projected images with various distorted degrees. We further develop a masked reprojection module across 2D and 3D latent space to bridge the gap between pre-trained 2D backbone and 3D point-wise features. Both quantitative and qualitative experimental re- sults on S3DIS and ScanNet V2 demonstrate the effective- ness of our method.11Code will be available here. Haowen Sun 0004, Yueqi Duan, Juncheng Yan, Jiwen Lu |
CVPR | 1 |
| 2024 | Make-Your-3D: Fast and Consistent Subject-Driven 3D Content Generation
Fangfu Liu, Hanyang Wang 0003, Haowen Sun 0004, Yueqi Duan |
ECCV (84) | 4 |
| 2024 | Learning Cross-Attention Point Transformer With Global Porous SamplingabstractIn this paper, we propose a point-based cross-attention transformer named CrossPoints with parametric Global Porous Sampling (GPS) strategy. The attention module is crucial to capture the correlations between different tokens for transformers. Most existing point-based transformers design multi-scale self-attention operations with down-sampled point clouds by the widely-used Farthest Point Sampling (FPS) strategy. However, FPS only generates sub-clouds with holistic structures, which fails to fully exploit the flexibility of points to generate diversified tokens for the attention module. To address this, we design a cross-attention module with parametric GPS and Complementary GPS (C-GPS) strategies to generate series of diversified tokens through controllable parameters. We show that FPS is a degenerated case of GPS, and the network learns more abundant relational information of the structure and geometry when we perform consecutive cross-attention over the tokens generated by GPS as well as C-GPS sampled points. More specifically, we set evenly-sampled points as queries and design our cross-attention layers with GPS and C-GPS sampled points as keys and values. In order to further improve the diversity of tokens, we design a deformable operation over points to adaptively adjust the points according to the input. Extensive experimental results on both shape classification and indoor scene segmentation tasks indicate promising boosts over the recent point cloud transformers. We also conduct ablation studies to show the effectiveness of our proposed cross-attention module with GPS strategy. Yueqi Duan, Haowen Sun 0004, Juncheng Yan, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Video Saliency Forecasting TransformerabstractVideo saliency prediction (VSP) aims to imitate eye fixations of humans. However, the potential of this task has not been fully exploited since existing VSP methods only focus on modeling visual saliency of the input previous frames. In this paper, we present the first attempt to extend this task to video saliency forecasting (VSF) by forecasting attention regions of consecutive future frames. To tackle this problem, we propose a video saliency forecasting transformer (VSFT) network built on a new encoder-decoder architecture. Different from existing VSP methods, our VSFT is the first pure-transformer based architecture in the VSP field and is freed from the dependency of the pretrained S3D model. In VSFT, the attention mechanism is exploited to capture spatial-temporal dependencies between the input past frames and the target future frame. We propose cross-attention guidance blocks (CAGB) to aggregate multi-level representation features to provide sufficient guidance for forecasting. We conduct comprehensive experiments on two benchmark datasets, DHF1K and Hollywoods-2. We investigate the saliency forecasting and predicting abilities of existing VSP methods by modifying the supervision signals. Experimental results demonstrate that our method achieves superior performance on both VSF and VSP tasks. Haowen Sun 0004, Yongming Rao, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |