Junao Shen

dblp:354/9745 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-2210-6503ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SeCuRe: Toward sparse-view 3D curve reconstruction with Gaussian Splatting
Tian Feng 0001, Haojie Dong, Jinkang Ji, Junao Shen, Feiyi Fan
Comput. Graph.4
2026 CAStGS: Context-aware scene stylization with 3D Gaussian splatting
Junao Shen, Ruihong Ye, Tian Feng 0001
Comput. Graph.2
2025 In2NeCT: Inter-class and Intra-class Neural Collapse Tuning for Semantic Segmentation of Imbalanced Remote Sensing Images
abstract
Remote sensing images (RSIs) are frequently characterized by multi-scale inter-class objects and inconsistently distributed objects due to scene limitations, which would cause a significant data imbalance challenging the corresponding semantic segmentation. Recent methods have leveraged various deep learning techniques to capture high-quality representations for RSI semantic segmentation, but are hardly capable of addressing the afore-mentioned challenge given their limited explorations towards the mechanisms behind the representations. The recently discovered Neural Collapse (NC) phenomenon in computer vision models suggests the simplex equiangular tight frame (ETF) as the optimal representation structure, which has motivated us to observe that the optimal structure of last-layer representations is disrupted and inter-class representations for minor classes tend to become closer to each other beacuse of data imbalance. To address these issues, we propose Inter-class and Intra-class Neural Collapse Tuning (In2NeCT) to optimize the representations that satisfy the simplex ETF, which facilitates the discrimination of inter-class representations and the coherence of intra-class representations. Extensive experiments on three datasets demonstrate that our In2NeCT consistently leads to significant improvements in performance and outperforms the state-of-the-art methods.
Junao Shen, Qiyun Hu, Tian Feng 0001, Xinyu Wang 0036, Hui Cui 0002, Sensen Wu, Wei Zhang 0243
AAAI1
2025 DRoLaS: Diffusion-Based Coarse-to-Fine Conditional Synthesis of Hierarchical Road Layouts
Shenao Dong, Bo Li 0173, Junao Shen, Tian Feng 0001
ICMR5
2025 WETR: Wireframe parsing using deformable Transformers
Jingwen Cui, Jinkang Ji, Junao Shen, Tianai Shen, Bo Li 0173, Tian Feng 0001
Comput. Graph.3
2025 CaRoLS: Condition-adaptive multi-level road layout synthesis
Tian Feng 0001, Bo Li 0173, Junao Shen
Comput. Graph.5
2025 SuraGS: Toward efficient few-shot novel view synthesis via surface-aware Gaussian splatting
Junao Shen, Tian Feng 0001, Haojie Dong, Jinkang Ji, Xinyu Wang 0036, Tianjia Shao
Comput. Graph.1
2024 CGMGM: A Cross-Gaussian Mixture Generative Model for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment unseen objects in a query image using a few pixel-wise annotated support images, thus expanding the capabilities of semantic segmentation. The main challenge lies in extracting sufficient information from the limited support images to guide the segmentation process. Conventional methods typically address this problem by generating single or multiple prototypes from the support images and calculating their cosine similarity to the query image. However, these methods often fail to capture meaningful information for modeling the de facto joint distribution of pixel and category. Consequently, they result in incomplete segmentation of foreground objects and mis-segmentation of the complex background. To overcome this issue, we propose the Cross Gaussian Mixture Generative Model (CGMGM), a novel Gaussian Mixture Models~(GMMs)-based FSS method, which establishes the joint distribution of pixel and category in both the support and query images. Specifically, our method initially matches the feature representations of the query image with those of the support images to generate and refine an initial segmentation mask. It then employs GMMs to accurately model the joint distribution of foreground and background using the support masks and the initial segmentation mask. Subsequently, a parametric decoder utilizes the posterior probability of pixels in the query image, by applying the Bayesian theorem, to the joint distribution, to generate the final segmentation mask. Experimental results on PASCAL-5i and COCO-20i datasets demonstrate our CGMGM's effectiveness and superior performance compared to the state-of-the-art methods.
Junao Shen, Kun Kuang 0001, Xinyu Wang 0036, Tian Feng 0001, Wei Zhang 0243
AAAI1
2024 SESAME: Toward Medical Image Segmentation via Foundation Model-assisted Semi-supervised Learning
abstract
Medical image segmentation is essential for diagnosis but requires expensive and time-consuming labeled data. Semi-supervised learning (SSL) mitigates this issue by using unlabeled data to improve generalization. However, current SSL methods encounter issues with inadaptive perturbations and low-quality pseudo-labels. Vision foundation models, such as SAM, have shown promise in segmentation. We propose SESAME, an SSL method integrating SAM and U-Net to improve labeling accuracy through a foundation model-assisted pipeline. In particular, we introduce a reliability score to address low-quality pseudo-labels and employ strategies for utilizing both reliable and unreliable images. Reliable images are associated with refined pseudo-labels via a conflict resolving strategy, whereas unreliable ones undergo a mutual region swapping strategy. Extensive experiments demonstrate that our SESAME outperforms representative methods for medical image segmentation.
Qiyun Hu, Junao Shen, Jinkang Ji, Xinyu Wang 0036, Tian Feng 0001, Hui Cui 0002
BIBM2
2024 Retinal Vessel Segmentation via Cross-attention Feature Fusion
abstract
Retinal vessel segmentation from fundus images is of significant importance for detecting and diagnosing common ocular diseases. Conventional deep learning-based methods for retinal vessel segmentation follow the U-Net framework with an encoder-decoder architecture and employ skip connections for the recovery of spatial information lost during downsampling. However, skip connections cannot consistently have positive contributions to segmentation performance, which is caused by the semantic incompatibility between encoder features and decoder features. Based on this observation, we propose CaFFNet, a Cross-attention Feature Fusion Network designed specifically for retinal vessel segmentation. Specifically, we improve skip connections by introducing a Cross-attention Feature Fusion (CaFF) module, which effectively mitigates the semantic gap between encoder and decoder feature maps by leveraging the cross-attention mechanism for feature fusion. Besides, we introduce a Dual-Branch Pooling Fusion (DBPF) module to address the loss of vessel spatial information during pooling and capture contextual details more effectively, so as to improve segmentation performance. Experimental results on three fundus image datasets demonstrate that our CaFFNet outperforms current representative methods for retinal vessel segmentation.
Tian Feng 0001, Junao Shen, Qiangguo Jin, Xinyu Wang 0036
ICME3
2024 WirePAuS: Auxiliary-free Single-shot Wireframe Parsing
abstract
Wireframe parsing aims to identify vectorized line segments as pairs of endpoints from an image. Conventional methods usually require field-specific knowledge for manual introduction of auxiliary processes or auxiliary learning tasks towards satisfactory performances. Such pipelines are, however, characterized by high complexity, insignificant efficiency, and limited space for further performance improvement. To address these issues, we propose WirePAuS, a novel Wireframe Parser with an Auxiliary-free Single-shot pipeline, which requires no auxiliary processes or auxiliary learning tasks. This is based on its capability to generate appropriate prior information from a prior-informed feature extractor, which incorporates frequency-domain and Hough-domain prior information on line segments in the backbone. Meanwhile, we devise a structurally compact pipeline that enables the parser to directly predict the focal midpoints as line objects and exploit their displacements for the corresponding endpoints. Our end-to-end trainable WirePAuS is capable to capture rich structural details during single-shot inference. Extensive experiments suggest that the proposed method reaches significantly improved performances for wireframe parsing and outperforms a series of state-of-the-art methods.
Jinkang Ji, Junao Shen, Xinyu Wang 0036, Tian Feng 0001, Sensen Wu
ICME2
2024 MuMoSNet: 3D MRI-based Brain Tumor Segmentation via Multi-modal and Multi-scale Feature Fusion
abstract
MRI images contain multi-modal information, introducing complexity to brain tumor segmentation. Recent studies have incorporated the Transformer model, given its exceptional capability to model long-range dependence, into convolutional neural networks (CNNs) to address limited receptive fields. However, such a hybrid strategy often neglects the inherent multimodal characteristics of MRI images and lacks the capacity to capture modality-specific features. In this paper, we propose a multi-modal and multi-scale feature fusion network (MuMoSNet) for brain tumor segmentation from 3D MRI images. Specifically, our MuMoSNet introduces a parallel ME-Transformer encoder alongside the CNN-based encoder in 3D U-Net to separately extract modality-specific features. Besides, we devise a multi-feature fusion (MuFF) module to learn affinity relationships between cross-modality shared features and modality-specific features, maximizing the exploration of multi-modal information. Extensive experiments on both BraTS21 and BraTS20 datasets suggest that our MuMoSNet outperforms current representative methods for brain tumor segmentation.
Hui Cui 0002, Junao Shen, Xinyu Wang 0036, Tian Feng 0001
ICME4