Shoutong Luo

dblp:276/3217 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0003-2967-0916ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2024 LDCNet: Long-Distance Context Modeling for Large-Scale 3D Point Cloud Scene Semantic Segmentation
abstract
Large-scale point cloud semantic segmentation is a challenging task in 3D computer vision. A key challenge is how to resolve ambiguities arising from locally high inter-class similarity. In this study, we introduce a solution by modeling long-distance contextual information to understand the scene's overall layout. The context sensitivity of previous methods is typically constrained to small blocks(e.g. 2m x 2m) and cannot be directly extended to the entire scene. For this reason, we propose Long-Distance Context Modeling Network(LDCNet). Our key insight is that keypoints are enough for inferring the layout of a scene. Therefore, we represent the entire scene using keypoints along with local descriptors and model long-distance context on these keypoints. Finally, we propagate the long-distance context information from keypoints back to non-keypoints. This allows our method to model long-distance context effectively. We conducted experiments on six datasets, demonstrating that our approach can effectively mitigate ambiguities. Our method performs well on large, irregular objects and exhibits good generalization for typical scenarios.
Shoutong Luo, Zhengxing Sun, Yi Wang 0125, Yunhan Sun
ACM Multimedia1
2022 Multi-view 3D Reconstruction from Video with Transformer
abstract
Multi-view 3D reconstruction is the base for many other applications in computer vision. Video provides multi-view images and temporal information, which can help us better complete the reconstruction goal. Redundant information handling in video and multi-view feature extraction and fusion become the key issues in the shape prior extraction for reconstruction. In this paper, inspired by the recent great success in Transformer models, we propose a transformer-based 3D reconstruction network. We formulate the multi-view 3D reconstruction into three parts: frame encoder, fusion module, and shape decoder. We apply several special used tokens and perform the fusion progressively in the encoder phase, called patch-level progressive fusion module. These tokens describe which part of the object the frame should focus on and the local structural detail progressively. Then we further design a transformer fusion module to aggregate the structure information. Finally, multi-head attention is utilized to build the transformer-based decoder to reuse the shallow features from encoder. In experiments not only can ours method achieve competitive performance, but it also has low model complexity and computation cost.
Yijie Zhong 0001, Zhengxing Sun, Yunhan Sun, Shoutong Luo
ICIP4
2022 Learning Semantic Segmentation on Unlabeled Real-World Indoor Point Clouds via Synthetic Data
abstract
The data-hungry nature of deep learning and the high cost of annotating point-level labels for point clouds make it difficult to apply semantic segmentation methods to unlabeled real-world indoor scenes. Therefore, label-efficient point cloud segmentation has become a promising research topic. We noticed that the online housing design platforms can provide a large number of synthetic indoor 3D scenes, which are created with semantic labels. In this paper, we propose to learn semantic segmentation on synthetic point clouds and adapt the model for unlabeled real-world data. The main challenge is that directly using models trained on synthetic data for real-world data produces poor results due to the large domain gap between synthetic and real-world data. We design a point cloud style transfer network and a feature discrimination network to reduce the domain gap in both the input space and the feature space. Experiments show that our approach significantly improves the performance on real-world data for models learned from synthetic data.
Youcheng Song, Zhengxing Sun, Yunjie Wu, Yunhan Sun, Shoutong Luo, Qian Li 0014
ICPR5
2022 Active Patterns Perceived for Stochastic Video Prediction
abstract
Predicting future scenes based on historical frames is challenging, especially when it comes to the complex uncertainty in nature. We observe that there is a divergence between spatial-temporal variations of active patterns and non-active patterns in a video, where these patterns constitute visual content and the former ones implicate more violent movement. This divergence enables active patterns the higher potential to act with more severe future uncertainty. Meanwhile, the existence of non-active patterns provides an opportunity for machines to examine some underlying rules with a mutual constraint between non-active patterns and active patterns. In order to solve this divergence, we provide a method called active patterns-perceived stochastic video prediction (ASVP) which allows active patterns to be perceived by neural networks during training. Our method starts with separating active patterns along with non-active ones from a video. Then, both scene-based prediction and active pattern-perceived prediction are conducted to respectively capture the variations within the whole scene and active patterns. Specially for active pattern-perceived prediction, a conditional generative adversarial network (CGAN) is exploited to model active patterns as conditions, with a variational autoencoder (VAE) for predicting the complex dynamics of active patterns. Additionally, a mutual constraint is designed to improve the learning procedure for the network to better understand underlying interacting rules among these patterns. Extensive experiments are conducted on both KTH human action and BAIR action-free robot pushing datasets with comparison to state-of-the-art works. Experimental results demonstrate the competitive performance of the proposed method as we expected. The released code and models are at https://github.com/tolearnmuch/ASVP.
Yechao Xu, Zhengxing Sun, Qian Li 0014, Yunhan Sun, Shoutong Luo
ACM Multimedia5
2022 Category-Sensitive Incremental Learning for Image-Based 3D Shape Reconstruction
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun
MMM (1)3
2022 Resolution-switchable 3D Semantic Scene Completion
abstract
Abstract Semantic scene completion (SSC) aims to recover the complete geometric structure as well as the semantic segmentation results from partial observations. Previous works could only perform this task at a fixed resolution. To handle this problem, we propose a new method that can generate results at different resolutions without redesigning and retraining. The basic idea is to decouple the direct connection between resolution and network structure. To achieve this, we convert feature volume generated by SSC encoders into a resolution adaptive feature and decode this feature via point. We also design a resolution‐adapted point sampling strategy for testing and a category‐based point sampling strategy for training to further handle this problem. The encoder of our method can be replaced by existing SSC encoders. We can achieve better results at other resolutions while maintaining the same accuracy as the original resolution results. Code and data are available at https://github.com/lstcutong/ReS-SSC .
Shoutong Luo, Zhengxing Sun, Yunhan Sun, Yi Wang 0125
Comput. Graph. Forum1
2022 Video supervised for 3D reconstruction from single image
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun
Multim. Tools Appl.3
2022 Learning indoor point cloud semantic segmentation from image-level labels
Youcheng Song, Zhengxing Sun, Qian Li 0014, Yunjie Wu, Yunhan Sun, Shoutong Luo
Vis. Comput.6
2020 Stable Video Style Transfer Based on Partial Convolution with Depth-Aware Supervision
abstract
As a very important research issue in digital media art, neural learning based video style transfer has attracted more and more attention. A lot of recent works import optical flow method to original image style transfer framework to preserve frame-coherency and prevent flicker. However, these methods highly rely on paired video datasets of content video and stylized video, which are often difficult to obtain. Another limitation of existing methods is that while maintaining inter-frame coherency, they will introduce strong ghosting artifacts. In order to address these problems, this paper has following contributions: (1).presents a novel training framework for video style transfer without dependency on video dataset of target style; (2).firstly focuses on the ghosting problem existing in most previous works and uses partial convolution-based strategy to utilize inter-frame context and correlation, together with additional depth loss as a constrain to the generated frames to suppress ghosting artifacts and preserve stability at the same time. Extensive experiments demonstrate that our method can produce natural and stable video frames with target style. Qualitative and quantitative comparisons also show that the proposed approach outperforms previous works in terms of overall image quality and inter-frame stability. To facilitate future research, we publish our experiment code at \urlhttps://github.com/Huage001/Artistic-Video-Partial-Conv-Depth-Loss.
Songhua Liu, Shoutong Luo, Zhengxing Sun
ACM Multimedia3