VLDB 2026 Research / reviewers in the wild / expert
Mengqi Rong
dblp:291/9315
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-4852-2345ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MC-MVSNet: When multi-view stereo meets monocular cues
Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen |
Pattern Recognit. | 2 |
| 2025 | Uncertainty Aware Multiple View Stereo Network with Accurate SupervisionabstractLearning-based multiple view stereo has gained significant attention recently. However, most methods rely on direct network supervision using provided ground-truth depth, which poses three inherent problems: resolution-dependent ground-truth artifacts, excessively challenging training examples (with relatively featureless textures), and use of less-viewed reference pixels for supervision, all of which hinder network optimization. To alleviate these problems, we propose an accurate network supervision paradigm that includes a ground-truth mask, an entropy mask, and a consistency mask, which provide more accurate supervision signals to aid network optimization. Furthermore, we introduce UANet, an uncertainty aware multi-view stereo network, which adaptively determines a pixel-wise search range using a dynamic range sampler (DRS) built upon estimation confidence and learned uncertainty. Experimental results on recent MVS datasets demonstrate the effectiveness of our method. Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen |
Comput. Vis. Media | 2 |
| 2024 | Revisiting Global Translation Estimation with Feature TracksabstractGlobal translation estimation is a highly challenging step in the global structure from motion (SfM) algorithm. Many existing methods rely solely on relative translations, leading to inaccuracies in low parallax scenes and degradation under collinear camera motion. While recent approaches aim to address these issues by incorporating feature tracks into objective functions, they are often sensitive to outliers. In this paper, we first revisit global translation estimation methods with feature tracks and categorize them into explicit and implicit methods. Then, we highlight the superiority of the objective function based on the cross-product distance metric and propose a novel explicit global translation estimation framework that integrates both relative translations and feature tracks as input. To enhance the accuracy of input observations, we re-estimate relative translations with the coplanarity constraint of the epipolar plane and propose a simple yet effective strategy to select reliable feature tracks. Finally, we demonstrate the effectiveness of our approach through experiments on urban image sequences and unordered Internet images, showcasing its superior accuracy and robustness compared to many state-of-the-art techniques. Peilin Tao, Hainan Cui, Mengqi Rong, Shuhan Shen |
CVPR | 3 |
| 2023 | 3D Semantic Segmentation of Aerial Photogrammetry Models Based on Orthographic ProjectionabstractSemantic segmentation of 3D scenes is one of the most important tasks in the field of computer vision and has attracted much attention. In this paper, we propose a novel framework for 3D semantic segmentation of aerial photogrammetry models, which uses orthographic projection to improve efficiency while still ensuring high precision, and can also be applied to multiple types of models (i.e., textured mesh or colored point cloud). In our pipeline, we first obtain RGB images and elevation images from the 3D scene through orthographic projection, then use the image semantic segmentation network to segment these images to obtain pixel-wise semantic predictions, and finally back-project the segmentation results to the 3D model for fusion. Specifically, for the image semantic segmentation model, we design a cross-modality feature aggregation module and a context guidance module based on category features, which assist the network in learning more discriminative features between different objects. For the 2D-3D semantic fusion, we combine the segmentation results of the 2D images with the geometric consistency of the 3D models for joint optimization, which further improves the accuracy of the 3D semantic segmentation. Extensive experiments on two large-scale urban scenes demonstrate the efficiency and feasibility of our algorithm and surpass the current mainstream 3D deep learning methods. Mengqi Rong, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Efficient 3D Scene Semantic Segmentation via Active Learning on Rendered 2D ImagesabstractInspired by Active Learning and 2D-3D semantic fusion, we proposed a novel framework for 3D scene semantic segmentation based on rendered 2D images, which could efficiently achieve semantic segmentation of any large-scale 3D scene with only a few 2D image annotations. In our framework, we first render perspective images at certain positions in the 3D scene. Then we continuously fine-tune a pre-trained network for image semantic segmentation and project all dense predictions to the 3D model for fusion. In each iteration, we evaluate the 3D semantic model and re-render images in several representative areas where the 3D segmentation is not stable and send them to the network for training after annotation. Through this iterative process of rendering-segmentation-fusion, it can effectively generate difficult-to-segment image samples in the scene, while avoiding complex 3D annotations, so as to achieve label-efficient 3D scene segmentation. Experiments on three large-scale indoor and outdoor 3D datasets demonstrate the effectiveness of the proposed method compared with other state-of-the-art. Mengqi Rong, Hainan Cui, Shuhan Shen |
IEEE Trans. Image Process. | 1 |
| 2022 | Active Learning Based 3D Semantic Labeling From Images and Videosabstract3D semantic segmentation is one of the most fundamental problems for 3D scene understanding and has attracted much attention in the field of computer vision. In this paper, we propose an active learning based 3D semantic labeling method for large-scale 3D mesh model generated from images or videos. Taking as input a 3D mesh model reconstructed from the image based 3D modeling system, coupled with the calibrated images, our method outputs a fine 3D semantic mesh model in which each facet is assigned a semantic label. There are three major steps in our framework: 2D semantic segmentation, 2D-3D semantic fusion, and batch image selection. A limited annotation image set is first used to fine-tune a pre-trained semantic segmentation network for obtaining the pixel-wise semantic probability maps. Then all these maps are back-projected into 3D space and fused on the 3D mesh model using Markov Random Field optimization, thus yield a preliminary 3D semantic mesh model and a heat model showing each facet’s confidence. This 3D semantic model is used as a reliable supervisor to select the parts that are not well segmented for manual annotation to boost the performance of the 2D semantic segmentation network, as well as the 3D mesh labeling, in the next iteration. This Training-Fusion-Selection process continues until the label assignment of the 3D mesh model becomes steady. By this means, we significantly reduce the amount for annotation but not the labeling quality of 3D semantic models. Extensive experiments demonstrate the effectiveness and generalization ability of our method on a wide variety of datasets. Mengqi Rong, Hainan Cui, Zhanyi Hu, Hanqing Jiang, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | 3D Semantic Labeling of Photogrammetry Meshes Based on Active LearningabstractAs different urban scenes are similar but still not completely consistent, coupled with the complexity of labeling directly in 3D, high-level understanding of 3D scenes has always been a tricky problem. In this paper, we propose a procedural approach for 3D semantic expression of urban scenes based on active learning. We first start with a small labeled image set to fine-tune a semantic segmentation network and then project its probability map onto a 3D mesh model for fusion, finally outputs a 3D semantic mesh model in which each facet has a semantic label and a heat model showing each facet's confidence. Our key observation is that our algorithm is iterative, in each iteration, we use the output semantic model as a supervision to select several valuable images for annotation to co-participate in the fine-tuning for overall improvement. In this way, we reduce the workload of labeling but not the quality of 3D semantic model. Using urban areas from two different cities, we show the potential of our method and demonstrate its effectiveness. Mengqi Rong, Shuhan Shen, Zhanyi Hu |
ICPR | 1 |