Yunhan Sun

dblp:218/6796 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
18since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Text Prompted Spatiotemporal Sequence Prediction with Text-Vision Prompt Refiner and Masked Diffusion Transformers
abstract
Classical spatiotemporal sequence prediction tasks are designed to forecast future image sequences based on historical observations. However, the inherent unpredictability of future events often renders this process uncontrollable due to infinite possibilities in nature, limiting broader applicability of this technology. In this study, we explore the utilization of text prompts to constrain probabilistic space of future outcomes, resulting more controllable future prediction complying with user intent. We primarily address two critical challenges in this research setting: (i) text-vision misalignment, where embeddings extracted by text pre-trained models are not strictly aligned with visual embeddings, leading to predictions semantically irrelevant to text prompts. (ii) Spatiotemporal modeling distortion, where the fixed observation interval during training causes the model to produce unrealistic results when reasoning longer time dimensions. To tackle these issues, we propose a text-prompted spatiotemporal sequence prediction (TPS2P) model, leveraging historical observations and textual prompts to predict probabilistic future outcomes. In this model, a text-vision prompt refiner (TV-Refiner) is introduced to provide aligned textual and historical visual embeddings for integrating the denoising diffusion prediction process. Additionally, a spatiotemporal-masked diffusion transformer (StMDiT) is proposed by exploiting masked attention in constituting spatial and temporal self-attention modules within latent diffusion processes, enabling the model to observe more sequences of varying spatiotemporal patterns during training. We conduct extensive experiments on Something-Something V2 (Sthv2) and BridgeData datasets. Reported results demonstrate that our TPS2P predicts more accurate and high-quality future sequences, more user-intent compliant by textual controllability.
Yechao Xu, Zhengxing Sun, Qian Li 0014, Yunhan Sun
ACM Multimedia4
2025 Spatiotemporal semantic structural representation learning for image sequence prediction
Yechao Xu, Zhengxing Sun, Yunhan Sun
Neurocomputing4
2025 Extended Receptive Field UDA Semantic Segmentation Based on Spatial Alignment and Knowledge Distillation
abstract
In recent years, unsupervised domain adaptation (UDA) has significantly advanced, addressing the issue of requiring large amounts of labeled data in deep learning. Some UDA strategies have effectively alleviated domain shift, but they are still struggling to tackle challenges such as the model’s inability to continuously extract features, the difficulty to achieve better segmentation boundaries in the target domain, and tend to neglect previously acquired knowledge. To address these issues, we propose an extended receptive field UDA semantic segmentation based on spatial alignment and knowledge distillation (ERF). Firstly, based on the idea of combining serial and parallel, we design a novel large continuous receptive field decoder (largeCF) to extract large continuous receptive field features. This approach alleviates the bias of model in feature extraction between objects of different sizes and simultaneously reduces model complexity. Secondly, we propose an edge consistency strategy that aligns edge features and matches the spatial arrangement between predicted and ground truth labels, improving edges segmentation accuracy of the target domain. Finally, we employ a knowledge distillation module to achieve an optimized student-teacher framework, where the teacher effectively guides the student to retain previously learned information, resulting in more accurate segmentation of the target domain. Experimental results demonstrated the effectiveness of the proposed approach, which achieved mIoU of 76.5% and 68.8% on UDA benchmark tasks GTA$\rightarrow $CityScapes and SYNTHIA$\rightarrow $CityScapes, respectively. The code is available at:https://github.com/fz-ss/ERF. Note to Practitioners—This paper focuses on the challenges of semantic segmentation in autonomous driving, particularly facing the issue of extensive manual annotation required for dense semantic labels. We propose a novel extended receptive field UDA semantic segmentation based on spatial alignment and knowledge distillation. The article begins by outlining the initial implementation process of UDA, laying the foundation for subsequent in-depth discussions. Subsequently, we conduct a theoretical analysis of the Large Continuous Decoder, Boundary Consistency Strategy, and Knowledge Distillation Scheme, which constitute the core components of our method. Finally, experiments on two UDA benchmark tasks demonstrate the feasibility of our approach. However, the performance gap between UDA and supervised semantic segmentation still exists. In future research, we will focus on reducing the feature gap between different domains. Additionally, we will strive to fully use image features and distilled features to make greater progress, thereby driving advancements in semantic segmentation for autonomous driving.
Yunna Song, Caisheng Liu, Suqin Bai, Xin Shu 0001, Yunhan Sun
IEEE Trans Autom. Sci. Eng.9
2025 IOFusion: instance segmentation and optical-flow guided 3D reconstruction in dynamic scenes
Haowei Zhu, Suqin Bai, Chenggen Wang, Yunhan Sun, Shucheng Huang
Vis. Comput.5
2024 LDCNet: Long-Distance Context Modeling for Large-Scale 3D Point Cloud Scene Semantic Segmentation
abstract
Large-scale point cloud semantic segmentation is a challenging task in 3D computer vision. A key challenge is how to resolve ambiguities arising from locally high inter-class similarity. In this study, we introduce a solution by modeling long-distance contextual information to understand the scene's overall layout. The context sensitivity of previous methods is typically constrained to small blocks(e.g. 2m x 2m) and cannot be directly extended to the entire scene. For this reason, we propose Long-Distance Context Modeling Network(LDCNet). Our key insight is that keypoints are enough for inferring the layout of a scene. Therefore, we represent the entire scene using keypoints along with local descriptors and model long-distance context on these keypoints. Finally, we propagate the long-distance context information from keypoints back to non-keypoints. This allows our method to model long-distance context effectively. We conducted experiments on six datasets, demonstrating that our approach can effectively mitigate ambiguities. Our method performs well on large, irregular objects and exhibits good generalization for typical scenarios.
Shoutong Luo, Zhengxing Sun, Yi Wang 0125, Yunhan Sun
ACM Multimedia4
2024 Transformer framework for depth-assisted UDA semantic segmentation
Yunna Song, Danping Zou, Caisheng Liu, Suqin Bai, Yunhan Sun
Eng. Appl. Artif. Intell.10
2024 Clear-Plenoxels: Floaters free radiance fields without neural networks
Weichen Yang, Suqin Bai, Zhen Ou, Yunhan Sun
Knowl. Based Syst.8
2024 Context-aware adaptive network for UDA semantic segmentation
Yunna Song, Zhen Ou, YueCheng Yu, Yunhan Sun
Multim. Syst.10
2024 EPM-Net: Efficient Feature Extraction, Point-Pair Feature Matching for Robust 6-D Pose Estimation
abstract
Estimating the 6-D poses of objects from RGB-D images holds great potential for several applications. However, given that the 6-D pose estimation accuracy is significantly affected by occlusion and noise between the objects in an image, this paper proposes a novel 6-D pose estimation method based on Efficient feature extraction and Point-pair feature matching. Specifically, we develop the Efficient channel attention Convolutional Neural Network (ECNN) and SO(3)-Encoder modules to extract 2-D features from the RGB image and SO(3)-equivariant features from the depth image, respectively. These features are fused in the DenseFusion module to obtain 3-D features in the camera space. Meanwhile, we exploit CAD model priors to obtain 3-D features in the model space through the model feature encoder, and then we globally regress the 3-D features in the camera and model space. According to these features, we generate oriented point clouds in each space, and then conduct point-pair feature matching to obtain pose information. Finally, we perform direct pose regression on the 3-D features in the camera and model space, and then resulting point-pair feature matching pose information is combined with the direct point-wise pose regression information to enhance pose prediction accuracy. Experimental results on three widely used benchmarking datasets demonstrate that our method achieves state-of-the-art performance, particularly for severe occluded scenes.
Danping Zou, Xin Shu 0001, Suqin Bai, Haowei Zhu, Yunhan Sun
IEEE Trans. Multim.9
2022 Multi-view 3D Reconstruction from Video with Transformer
abstract
Multi-view 3D reconstruction is the base for many other applications in computer vision. Video provides multi-view images and temporal information, which can help us better complete the reconstruction goal. Redundant information handling in video and multi-view feature extraction and fusion become the key issues in the shape prior extraction for reconstruction. In this paper, inspired by the recent great success in Transformer models, we propose a transformer-based 3D reconstruction network. We formulate the multi-view 3D reconstruction into three parts: frame encoder, fusion module, and shape decoder. We apply several special used tokens and perform the fusion progressively in the encoder phase, called patch-level progressive fusion module. These tokens describe which part of the object the frame should focus on and the local structural detail progressively. Then we further design a transformer fusion module to aggregate the structure information. Finally, multi-head attention is utilized to build the transformer-based decoder to reuse the shallow features from encoder. In experiments not only can ours method achieve competitive performance, but it also has low model complexity and computation cost.
Yijie Zhong 0001, Zhengxing Sun, Yunhan Sun, Shoutong Luo
ICIP3
2022 Learning Semantic Segmentation on Unlabeled Real-World Indoor Point Clouds via Synthetic Data
abstract
The data-hungry nature of deep learning and the high cost of annotating point-level labels for point clouds make it difficult to apply semantic segmentation methods to unlabeled real-world indoor scenes. Therefore, label-efficient point cloud segmentation has become a promising research topic. We noticed that the online housing design platforms can provide a large number of synthetic indoor 3D scenes, which are created with semantic labels. In this paper, we propose to learn semantic segmentation on synthetic point clouds and adapt the model for unlabeled real-world data. The main challenge is that directly using models trained on synthetic data for real-world data produces poor results due to the large domain gap between synthetic and real-world data. We design a point cloud style transfer network and a feature discrimination network to reduce the domain gap in both the input space and the feature space. Experiments show that our approach significantly improves the performance on real-world data for models learned from synthetic data.
Youcheng Song, Zhengxing Sun, Yunjie Wu, Yunhan Sun, Shoutong Luo, Qian Li 0014
ICPR4
2022 Active Patterns Perceived for Stochastic Video Prediction
abstract
Predicting future scenes based on historical frames is challenging, especially when it comes to the complex uncertainty in nature. We observe that there is a divergence between spatial-temporal variations of active patterns and non-active patterns in a video, where these patterns constitute visual content and the former ones implicate more violent movement. This divergence enables active patterns the higher potential to act with more severe future uncertainty. Meanwhile, the existence of non-active patterns provides an opportunity for machines to examine some underlying rules with a mutual constraint between non-active patterns and active patterns. In order to solve this divergence, we provide a method called active patterns-perceived stochastic video prediction (ASVP) which allows active patterns to be perceived by neural networks during training. Our method starts with separating active patterns along with non-active ones from a video. Then, both scene-based prediction and active pattern-perceived prediction are conducted to respectively capture the variations within the whole scene and active patterns. Specially for active pattern-perceived prediction, a conditional generative adversarial network (CGAN) is exploited to model active patterns as conditions, with a variational autoencoder (VAE) for predicting the complex dynamics of active patterns. Additionally, a mutual constraint is designed to improve the learning procedure for the network to better understand underlying interacting rules among these patterns. Extensive experiments are conducted on both KTH human action and BAIR action-free robot pushing datasets with comparison to state-of-the-art works. Experimental results demonstrate the competitive performance of the proposed method as we expected. The released code and models are at https://github.com/tolearnmuch/ASVP.
Yechao Xu, Zhengxing Sun, Qian Li 0014, Yunhan Sun, Shoutong Luo
ACM Multimedia4
2022 Category-Sensitive Incremental Learning for Image-Based 3D Shape Reconstruction
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun
MMM (1)4
2022 Resolution-switchable 3D Semantic Scene Completion
abstract
Abstract Semantic scene completion (SSC) aims to recover the complete geometric structure as well as the semantic segmentation results from partial observations. Previous works could only perform this task at a fixed resolution. To handle this problem, we propose a new method that can generate results at different resolutions without redesigning and retraining. The basic idea is to decouple the direct connection between resolution and network structure. To achieve this, we convert feature volume generated by SSC encoders into a resolution adaptive feature and decode this feature via point. We also design a resolution‐adapted point sampling strategy for testing and a category‐based point sampling strategy for training to further handle this problem. The encoder of our method can be replaced by existing SSC encoders. We can achieve better results at other resolutions while maintaining the same accuracy as the original resolution results. Code and data are available at https://github.com/lstcutong/ReS-SSC .
Shoutong Luo, Zhengxing Sun, Yunhan Sun, Yi Wang 0125
Comput. Graph. Forum3
2022 Video supervised for 3D reconstruction from single image
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun
Multim. Tools Appl.4
2022 Learning indoor point cloud semantic segmentation from image-level labels
Youcheng Song, Zhengxing Sun, Qian Li 0014, Yunjie Wu, Yunhan Sun, Shoutong Luo
Vis. Comput.5
2021 Shape-Pose Ambiguity in Learning 3D Reconstruction from Images
Yunjie Wu, Zhengxing Sun, Youcheng Song, Yunhan Sun, Yijie Zhong 0001
AAAI4
2021 A self-supervised method of single-image depth estimation by feeding forward information using max-pooling layers
Yunhan Sun, Suqin Bai, Zhengxing Sun, Zhaohui Tian
Vis. Comput.2
2020 Slicenet: Slice-Wise 3D Shapes Reconstruction from Single Image
abstract
3D object reconstruction from a single image is a highly ill-posed problem, requiring strong prior knowledge of 3D shapes. Deep learning methods are popular for this task. Especially, most works utilized 3D deconvolution to generate 3D shapes. However, the resolution of results is limited by the high resource consumption of 3D deconvolution. In this paper, we propose SliceNet, sequentially generating 2D slices of 3D shapes with shared 2D deconvolution parameters. To capture relations between slices, the RNN is also introduced. Our model has three main advantages: First, the introduction of RNN allows the CNN to focus more on local geometry details,improving the results’ fine-grained plausibility. Second, replacing 3D deconvolution with 2D deconvolution reducs much consumption of memory, enabling higher resolution of final results. Third, an slice-aware attention mechanism is designed to provide dynamic information for each slice’s generation, which helps modeling the difference between multiple slices, making the learning process easier. Experiments on both synthesized data and real data illustrate the effectiveness of our method.
Yunjie Wu, Zhengxing Sun, Youcheng Song, Yunhan Sun
ICASSP4
2020 Single View Depth Estimation via Dense Convolution Network with Self-supervision
Yunhan Sun, Suqin Bai, Zhengxing Sun
MMM (2)1
2020 Progressive decomposition: a method of coarse-to-fine image parsing using stacked networks
Yunhan Sun, Jiagao Hu, Zhengxing Sun
Multim. Tools Appl.1
2019 Detecting Robust Co-Saliency with Recurrent Co-Attention Neural Network
abstract
Effective feature representations which should not only express the images individual properties, but also reflect the interaction among group images are essentially crucial for robust co-saliency detection. This paper proposes a novel deep learning co-saliency detection approach which simultaneously learns single image properties and robust group feature in a recurrent manner. Specifically, our network first extracts the semantic features of each image. Then, a specially designed Recurrent Co-Attention Unit (RCAU) will explore all images in the group recurrently to generate the final group representation using the co-attention between images, and meanwhile suppresses noisy information. The group feature which contains complementary synergetic information is later merged with the single image features which express the unique properties to infer robust co-saliency. We also propose a novel co-perceptual loss to make full use of interactive relationships of whole images in the training group as the supervision in our end-to-end training process. Extensive experimental results demonstrate the superiority of our approach in comparison with the state-of-the-art methods.
Bo Li 0061, Zhengxing Sun, Lv Tang, Yunhan Sun
IJCAI4
2018 Progressive Refinement: A Method of Coarse-to-Fine Image Parsing Using Stacked Network
abstract
To parse images into fine-grained semantic parts, the complex fine-grained elements will put it in trouble when using off-the-shelf semantic segmentation networks. In this paper, for image parsing task, we propose to parse images from coarse to fine with progressively refined semantic classes. It is achieved by stacking the segmentation layers in a segmentation network several times. The former segmentation module parses images at a coarser-grained level, and the result will be feed to the following one to provide effective contextual clues for the finer-grained parsing. To recover the details of small structures, we add skip connections from shallow layers of the network to fine-grained parsing modules. As for the network training, we merge classes in groundtruth to get coarse-to-fine label maps, and train the stacked network with these hierarchical supervision end-to-end. Our coarse-to-fine stacked framework can be injected into many advanced neural networks to improve the parsing results. Extensive evaluations on several public datasets including face parsing and human parsing well demonstrate the superiority of our method.
Jiagao Hu, Zhengxing Sun, Yunhan Sun
ICME3
2018 Accumulative image categorization: a personal photo classification method for progressive collection
Jiagao Hu, Zhengxing Sun, Yunhan Sun
Multim. Tools Appl.3