Yujing Lou

dblp:230/8004 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-6292-8953ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
3D vision · 85% Segmentation and scene understanding · 6% Image recognition and object detection · 5%
Computer graphics and multimedia
2 papers
Rendering · 93% Geometric modeling and processing · 7%

Topics — the 22 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
point cloud analysis
1.732023
CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame · AAAI 2023
PRIN/SPRIN: On Extracting Point-Wise Rotation Invariant Features · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Pointwise Rotation-Invariant Network with Adaptive Sampling and 3D Spherical Voxel Convolution · AAAI 2020
Computer vision › 3D vision › point cloud processing
rotation-invariant feature extraction
1.022022
PRIN/SPRIN: On Extracting Point-Wise Rotation Invariant Features · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Pointwise Rotation-Invariant Network with Adaptive Sampling and 3D Spherical Voxel Convolution · AAAI 2020
Computer vision › 3D vision
3d scene understanding
0.912025
HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting · CVPR 2025
Rendering › gaussian splatting
3d gaussian splatting
0.912025
HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting · CVPR 2025
Rendering
novel view synthesis
0.912025
HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting · CVPR 2025
Computer vision › 3D vision › pose estimation
rotation estimation
0.712023
CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame · AAAI 2023
Computer vision › 3D vision › local feature descriptor
rotation-invariant descriptor
0.712023
CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame · AAAI 2023
Computer vision › 3D vision
3d object detection
0.612022
Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes · CVPR 2022
Computer vision › Image recognition and object detection › object detection
oriented object detection
0.612022
Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes · CVPR 2022
Computer vision › 3D vision › point cloud analysis
point cloud classification
0.612022
PRIN/SPRIN: On Extracting Point-Wise Rotation Invariant Features · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › 3D vision
point cloud processing
0.612022
Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes · CVPR 2022
Computer vision › Segmentation and scene understanding
semantic segmentation
0.612022
Understanding Pixel-Level 2D Image Semantics With 3D Keypoint Knowledge Engine · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › 3D vision › 3d object detection › proposal-based 3d object detection
vote-based 3d detection
0.612022
Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes · CVPR 2022
Computer vision › 3D vision › low-level vision › feature detection
keypoint detection
0.512021
Localization with Sampling-Argmax · NeurIPS 2021
Computer vision › 3D vision › 3d shape analysis
3d keypoint detection
0.412020
KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human Annotations · CVPR 2020
Computer vision › 3D vision › 3d shape analysis
3d shape understanding
0.412020
Human Correspondence Consensus for 3D Object Semantic Understanding · ECCV (22) 2020
Computer vision › 3D vision › point cloud analysis
point cloud classification and segmentation
0.412020
Pointwise Rotation-Invariant Network with Adaptive Sampling and 3D Spherical Voxel Convolution · AAAI 2020
Computer vision › 3D vision › correspondence estimation
semantic correspondence
0.412020
Human Correspondence Consensus for 3D Object Semantic Understanding · ECCV (22) 2020
Computer vision › Segmentation and scene understanding
part segmentation
0.212022
PRIN/SPRIN: On Extracting Point-Wise Rotation Invariant Features · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › 3D vision › feature matching › 3d correspondence
3d feature matching
0.112020
Pointwise Rotation-Invariant Network with Adaptive Sampling and 3D Spherical Voxel Convolution · AAAI 2020
Computer vision › 3D vision
3d shape analysis
0.112020
Human Correspondence Consensus for 3D Object Semantic Understanding · ECCV (22) 2020
Geometric modeling and processing
shape analysis
0.112020
KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human Annotations · CVPR 2020

Methods — techniques the papers use, named apart from their topics

multi-view supervision · 1.72d gaussian representation · 1.7spherical voxel convolution · 1.0point re-sampling · 1.0centrifugal reference frame · 0.7attention-based down-sampling · 0.7local canonical coordinates · 0.6canonical voting · 0.6back-projection checking · 0.63d keypoint projection · 0.6human annotation aggregation · 0.4fidelity loss minimization · 0.4
YearPublicationVenuePosition
2025 HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting
abstract
Generating high-quality novel view renderings of 3D Gaussian Splatting (3DGS) in scenes featuring transient objects is challenging. We propose a novel hybrid representation, termed as HybridGS, using 2D Gaussians for transient objects per image and maintaining traditional 3D Gaussians for the whole static scenes. 3DGS is suited for modeling static scenes that assume multi-view consistency, but the transient objects appear occasionally and do not adhere to the assumption, thus we model them as planar objects from a single view by 2D Gaussians. Our novel representation decomposes the scene from the perspective of fundamental viewpoint consistency. Additionally, we present a multi-view supervision method for 3DGS that leverages information from co-visible regions, further enhancing the distinctions between the transients and statics. Then, we propose a straightforward yet effective multi-stage training strategy to ensure robust training and view synthesis. Experiments on benchmarks show our state-of-the-art performance of novel view synthesis in indoor and outdoor scenes, even in the presence of distracting elements. Project page: https://gujiaqivadin.github.io/hybridgs/
Jiaqi Gu 0004, Lubin Fan, Bojian Wu, Yujing Lou, Renjie Chen 0001, Ligang Liu 0001, Jieping Ye
CVPR5
2025 SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
abstract
While vision language models (VLMs) excel in 2D semantic visual understanding, their ability to quantitatively reason about 3D spatial relationships remains underexplored due to the deficiency of spatial representation ability of 2D images. In this paper, we analyze the problem hindering VLMs’ spatial understanding abilities and propose SD-VLM, a novel framework that significantly enhances fundamental spatial perception abilities of VLMs through two key contributions: (1) propose Massive Spatial Measuring and Understanding (MSMU) dataset with precise spatial annotations, and (2) introduce a simple depth positional encoding method strengthening VLMs’ spatial awareness. MSMU dataset includes massive quantitative spatial tasks with 700K QA pairs, 2.5M physical numerical annotations, and 10K chain-of-thought augmented samples. We have trained SD-VLM, a strong generalist VLM which shows superior quantitative spatial measuring and understanding capability. SD-VLM not only achieves state-of-the-art performance on our proposed MSMU-Bench, but also shows spatial generalization abilities on other spatial understanding benchmarks including Q-Spatial and SpatialRGPTBench. Extensive experiments demonstrate that SD-VLM outperforms GPT-4o and Intern-VL3-78B by 26.91% and 25.56% respectively on MSMU-Bench. Code and models are released at https://github.com/cpystan/SD-VLM.
Pingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo, Lubin Fan, Lin Yang 0011, Lizhuang Ma, Jieping Ye
NeurIPS2
2023 CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame
abstract
Various recent methods attempt to implement rotation-invariant 3D deep learning by replacing the input coordinates of points with relative distances and angles. Due to the incompleteness of these low-level features, they have to undertake the expense of losing global information. In this paper, we propose the CRIN, namely Centrifugal Rotation-Invariant Network. CRIN directly takes the coordinates of points as input and transforms local points into rotation-invariant representations via centrifugal reference frames. Aided by centrifugal reference frames, each point corresponds to a discrete rotation so that the information of rotations can be implicitly stored in point features. Unfortunately, discrete points are far from describing the whole rotation space. We further introduce a continuous distribution for 3D rotations based on points. Furthermore, we propose an attention-based down-sampling strategy to sample points invariant to rotations. A relation module is adopted at last for reinforcing the long-range dependencies between sampled points and predicts the anchor point for unsupervised rotation estimation. Extensive experiments show that our method achieves rotation invariance, accurately estimates the object rotation, and obtains state-of-the-art results on rotation-augmented classification and part segmentation. Ablation studies validate the effectiveness of the network design.
Yujing Lou, Zelin Ye, Yang You 0004, Nianjuan Jiang, Jiangbo Lu, Lizhuang Ma, Cewu Lu
AAAI1
2022 Canonical Voting: Towards Robust Oriented Bounding Box Detection in 3D Scenes
abstract
3D object detection has attracted much attention thanks to the advances in sensors and deep learning methods for point clouds. Current state-of-the-art methods like VoteNet regress direct offset towards object centers and box orientations with an additional Multi-Layer-Perceptron network. Both their offset and orientation predictions are not accurate due to the fundamental difficulty in rotation classification. In the work, we disentangle the direct offset into Local Canonical Coordinates (LCC), box scales and box orientations. Only LCC and box scales are regressed, while box orientations are generated by a canonical voting scheme. Finally, an LCC-aware back-projection checking algorithm iteratively cuts out bounding boxes from the generated vote maps, with the elimination of false positives. Our model achieves state-of-the-art performance on three standard real-world benchmarks: ScanNet, SceneNN and SUN RGB-D. Our code is available on https://github.com/qq456cvb/CanonicalVoting.
Yang You 0004, Zelin Ye, Yujing Lou, Chengkun Li, Yong-Lu Li 0001, Lizhuang Ma, Cewu Lu
CVPR3
2022 Understanding Pixel-Level 2D Image Semantics With 3D Keypoint Knowledge Engine
abstract
Pixel-level 2D object semantic understanding is an important topic in computer vision and could help machine deeply understand objects (e.g., functionality and affordance) in our daily life. However, most previous methods directly train on correspondences in 2D images, which is end-to-end but loses plenty of information in 3D spaces. In this paper, we propose a new method on predicting image corresponding semantics in 3D domain and then projecting them back onto 2D images to achieve pixel-level understanding. In order to obtain reliable 3D semantic labels that are absent in current image datasets, we build a large scale keypoint knowledge engine called KeypointNet, which contains 103,450 keypoints and 8,234 3D models from 16 object categories. Our method leverages the advantages in 3D vision and can explicitly reason about objects self-occlusion and visibility. We show that our method gives comparative and even superior results on standard semantic benchmarks.
Yang You 0004, Chengkun Li, Yujing Lou, Zhoujun Cheng, Liangwei Li, Lizhuang Ma, Cewu Lu
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 PRIN/SPRIN: On Extracting Point-Wise Rotation Invariant Features
abstract
Point cloud analysis without pose priors is very challenging in real applications, as the orientations of point clouds are often unknown. In this paper, we propose a brand new point-set learning framework PRIN, namely, Point-wise Rotation Invariant Network, focusing on rotation invariant feature extraction in point clouds analysis. We construct spherical signals by Density Aware Adaptive Sampling to deal with distorted point distributions in spherical space. Spherical Voxel Convolution and Point Re-sampling are proposed to extract rotation invariant features for each point. In addition, we extend PRIN to a sparse version called SPRIN, which directly operates on sparse point clouds. Both PRIN and SPRIN can be applied to tasks ranging from object classification, part segmentation, to 3D feature matching and label alignment. Results show that, on the dataset with randomly rotated point clouds, SPRIN demonstrates better performance than state-of-the-art methods without any data augmentation. We also provide thorough theoretical proof and analysis for point-wise rotation invariance achieved by our methods. The code to reproduce our results will be made publicly available.
Yang You 0004, Yujing Lou, Ruoxi Shi, Yu-Wing Tai, Lizhuang Ma, Cewu Lu
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Localization with Sampling-Argmax
abstract
Soft-argmax operation is commonly adopted in detection-based methods to localize the target position in a differentiable manner. However, training the neural network with soft-argmax makes the shape of the probability map unconstrained. Consequently, the model lacks pixel-wise supervision through the map during training, leading to performance degradation. In this work, we propose sampling-argmax, a differentiable training method that imposes implicit constraints to the shape of the probability map by minimizing the expectation of the localization error. To approximate the expectation, we introduce a continuous formulation of the output distribution and develop a differentiable sampling process. The expectation can be approximated by calculating the average error of all samples drawn from the output distribution. We show that sampling-argmax can seamlessly replace the conventional soft-argmax operation on various localization tasks. Comprehensive experiments demonstrate the effectiveness and flexibility of the proposed method. Code is available at https://github.com/Jeff-sjtu/sampling-argmax
Ruiqi Shi, Yujing Lou, Yong-Lu Li 0001, Cewu Lu
NeurIPS4
2020 Pointwise Rotation-Invariant Network with Adaptive Sampling and 3D Spherical Voxel Convolution
abstract
Point cloud analysis without pose priors is very challenging in real applications, as the orientations of point clouds are often unknown. In this paper, we propose a brand new point-set learning framework PRIN, namely, Pointwise Rotation-Invariant Network, focusing on rotation-invariant feature extraction in point clouds analysis. We construct spherical signals by Density Aware Adaptive Sampling to deal with distorted point distributions in spherical space. In addition, we propose Spherical Voxel Convolution and Point Re-sampling to extract rotation-invariant features for each point. Our network can be applied to tasks ranging from object classification, part segmentation, to 3D feature matching and label alignment. We show that, on the dataset with randomly rotated point clouds, PRIN demonstrates better performance than state-of-the-art methods without any data augmentation. We also provide theoretical analysis for the rotation-invariance achieved by our methods.
Yang You 0004, Yujing Lou, Yu-Wing Tai, Lizhuang Ma, Cewu Lu
AAAI2
2020 KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human Annotations
abstract
Detecting 3D objects keypoints is ofgreat interest to the areas of both graphics and computer vision. There have been several 2D and 3D keypoint datasets aiming to address this problem in a data-driven way. These datasets, however, either lack scalability or bring ambiguity to the definition of keypoints. Therefore, we present KeypointNet: the first large-scale and diverse 3D keypoint dataset that contains 83,231 keypoints and 8,329 3D models from 16 object categories, by leveraging numerous human annotations. To handle the inconsistency between annotations from different people, we propose a novel method to aggregate these keypoints automatically, through minimization of a fidelity loss. Finally, ten state-of-the-art methods are benchmarked on our proposed dataset.
Yang You 0004, Yujing Lou, Chengkun Li, Zhoujun Cheng, Liangwei Li, Lizhuang Ma, Cewu Lu
CVPR2
2020 Human Correspondence Consensus for 3D Object Semantic Understanding
Yujing Lou, Yang You 0004, Chengkun Li, Zhoujun Cheng, Liangwei Li, Lizhuang Ma, Cewu Lu
ECCV (22)1