Yingshu Chen

dblp:298/7733 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-5418-2813ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
5 papers
Visual content generation and editing · 37% Rendering · 29% Computational photography and imaging · 29%
Artificial intelligence
5 papers
Video understanding and tracking · 49% Generative modeling · 18% 3D vision · 16%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › style transfer › 3d content stylization
3d scene stylization
1.622025
Advances in 3D Neural Stylization: A Survey · Int. J. Comput. Vis. 2025
StyleCity: Large-Scale 3D Urban Scenes Stylization · ECCV (59) 2024
Computer vision › Video understanding and tracking
object tracking
1.522025
360VOTS: Visual Object Tracking and Segmentation in Omnidirectional Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025
360VOT: A New Benchmark Dataset for Omnidirectional Visual Object Tracking · ICCV 2023
Machine learning › Generative modeling › style transfer
neural style transfer
0.912025
Advances in 3D Neural Stylization: A Survey · Int. J. Comput. Vis. 2025
Computer vision › Video understanding and tracking
video object segmentation
0.912025
360VOTS: Visual Object Tracking and Segmentation in Omnidirectional Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computational photography and imaging
camera calibration
0.912025
SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting · ICLR 2025
Rendering
gaussian splatting
0.912025
SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting · ICLR 2025
Computational photography and imaging
omnidirectional imaging
0.912025
SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting · ICLR 2025
Rendering › neural rendering
radiance field
0.912025
SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting · ICLR 2025
Computer vision › 3D vision
3d scene understanding
0.812024
StyleCity: Large-Scale 3D Urban Scenes Stylization · ECCV (59) 2024
Robotics › Robot navigation and mapping › mobile robot perception
omnidirectional perception
0.712023
360VOT: A New Benchmark Dataset for Omnidirectional Visual Object Tracking · ICCV 2023
Virtual and augmented reality › immersive video
360-degree video
0.312025
360VOTS: Visual Object Tracking and Segmentation in Omnidirectional Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Segmentation and scene understanding
scene understanding
0.212022
Neural Scene Decoration from a Single Photograph · ECCV (23) 2022

Methods — techniques the papers use, named apart from their topics

3d gaussian splatting · 2.4neural network · 1.7extended bounding field-of-view · 1.7neural style transfer · 1.5scene synthesis · 1.1neural rendering · 1.1differentiable camera model · 0.9equirectangular projection · 0.7bounding field-of-view · 0.7
YearPublicationVenuePosition
2025 SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting
abstract
360-degree cameras streamline data collection for radiance field 3D reconstruction by capturing comprehensive scene data. However, traditional radiance field methods do not address the specific challenges inherent to 360-degree images. We present SC-OmniGS, a novel self-calibrating omnidirectional Gaussian splatting system for fast and accurate omnidirectional radiance field reconstruction using 360-degree images. Rather than converting 360-degree images to cube maps and performing perspective image calibration, we treat 360-degree images as a whole sphere and derive a mathematical framework that enables direct omnidirectional camera pose calibration accompanied by 3D Gaussians optimization. Furthermore, we introduce a differentiable omnidirectional camera model in order to rectify the distortion of real-world data for performance enhancement. Overall, the omnidirectional camera intrinsic model, extrinsic poses, and 3D Gaussians are jointly optimized by minimizing weighted spherical photometric loss. Extensive experiments have demonstrated that our proposed SC-OmniGS is able to recover a high-quality radiance field from noisy camera poses or even no pose prior in challenging scenarios characterized by wide baselines and non-object-centric configurations. The noticeable performance gain in the real-world dataset captured by consumer-grade omnidirectional cameras verifies the effectiveness of our general omnidirectional camera model in reducing the distortion of 360-degree images.
Huajian Huang, Yingshu Chen, Tristan Braud, Sai-Kit Yeung
ICLR2
2025 Localized Gaussian Splatting Editing with Contextual Awareness
abstract
Recent advancements in text-guided 3D object generation using diffusion priors struggle with illumination inconsistencies when applied to scene editing tasks like object replacement or insertion. To address this, we propose an illumination-aware 3D scene editing pipeline for 3D Gaussian Splatting (3DGS). Our method leverages state-of-the-art 2D diffusion inpainting [56] to handle global illumination context effectively. Specifically, we identify representative anchor views that capture scene-wide illumination, inpaint them using 2D diffusion models, and integrate the results into a coarse-to-fine 3DGS optimization process. In the fine step, we introduce Depth-guided Inpainting Score Distillation Sampling (DI-SDS) to refine geometry and texture details, capitalizing on the diversity of 2D priors. Our approach achieves locally precise edits with globally consistent illumination, demonstrating robustness in real scenes with highlights and shadows. Comparisons show superior results over state-of-the-art text-to-3D editing methods. Project page: https://corneliushsiao.github.io/GSLE.html.
Hanyuan Xiao, Yingshu Chen, Huajian Huang, Haolin Xiong, Pratusha Prasad
WACV2
2025 Advances in 3D Neural Stylization: A Survey
abstract
Abstract Modern artificial intelligence offers a novel and transformative approach to creating digital art across diverse styles and modalities like images, videos and 3D data, unleashing the power of creativity and revolutionizing the way that we perceive and interact with visual content. This paper reports on recent advances in stylized 3D asset creation and manipulation with the expressive power of neural networks. We establish a taxonomy for neural stylization, considering crucial design choices such as scene representation, guidance data, optimization strategies, and output styles. Building on such taxonomy, our survey first revisits the background of neural stylization on 2D images, and then presents in-depth discussions on recent neural stylization methods for 3D data, accompanied by a benchmark evaluating selected mesh and neural field stylization methods. Based on the insights gained from the survey, we highlight the practical significance, open challenges, future research, and potential impacts of neural stylization, which facilitates researchers and practitioners to navigate the rapidly evolving landscape of 3D content creation using modern artificial intelligence.
Yingshu Chen, Guocheng Shao, Ka-Chun Shum, Binh-Son Hua, Sai-Kit Yeung
Int. J. Comput. Vis.1
2025 360VOTS: Visual Object Tracking and Segmentation in Omnidirectional Videos
abstract
Visual object tracking and segmentation in omnidirectional videos are challenging due to the wide field-of-view and large spherical distortion brought by 360$^{\circ }$∘ images. To alleviate these problems, we introduce a novel representation, extended bounding field-of-view (eBFoV), for target localization and use it as the foundation of a general 360 tracking framework which is applicable for both omnidirectional visual object tracking and segmentation tasks. Building upon our previous work on omnidirectional visual object tracking (360VOT), we propose a comprehensive dataset and benchmark that incorporates a new component called omnidirectional video object segmentation (360VOS). The 360VOS dataset includes 290 sequences accompanied by dense pixel-wise masks and covers a broader range of target categories. To support both the development and evaluation of algorithms in this domain, we divide the dataset into a training subset with 170 sequences and a testing subset with 120 sequences. Furthermore, we tailor evaluation metrics for both omnidirectional tracking and segmentation to ensure rigorous assessment. Through extensive experiments, we benchmark state-of-the-art approaches and demonstrate the effectiveness of our proposed 360 tracking framework and training dataset.
Yinzhe Xu, Huajian Huang, Yingshu Chen, Sai-Kit Yeung
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 StyleCity: Large-Scale 3D Urban Scenes Stylization
Yingshu Chen, Huajian Huang, Tuan-Anh Vu, Ka-Chun Shum, Sai-Kit Yeung
ECCV (59)1
2023 360VOT: A New Benchmark Dataset for Omnidirectional Visual Object Tracking
abstract
360° images can provide an omnidirectional field of view which is important for stable and long-term scene perception. In this paper, we explore 360° images for visual object tracking and perceive new challenges caused by large distortion, stitching artifacts, and other unique attributes of 360° images. To alleviate these problems, we take advantage of novel representations of target localization, i.e., bounding field-of-view, and then introduce a general 360 tracking framework that can adopt typical trackers for omnidirectional tracking. More importantly, we propose a new large-scale omnidirectional tracking benchmark dataset, 360VOT, in order to facilitate future research. 360VOT contains 120 sequences with up to 113K high-resolution frames in equirectangular projection. The tracking targets cover 32 categories in diverse scenarios. Moreover, we provide 4 types of unbiased ground truth, including (rotated) bounding boxes and (rotated) bounding field-of-views, as well as new metrics tailored for 360° images which allow for the accurate evaluation of omnidirectional tracking performance. Finally, we extensively evaluated 20 state-of-the-art visual trackers and provided a new baseline for future comparisons. Homepage: https://360vot.hkustvgd.com
Huajian Huang, Yinzhe Xu, Yingshu Chen, Sai-Kit Yeung
ICCV3
2023 Cross-Domain Autonomous Driving Perception Using Contrastive Appearance Adaptation
abstract
Addressing domain shifts for complex perception tasks in autonomous driving has long been a challenging problem. In this paper, we show that existing domain adaptation methods pay little attention to the content mismatch issue between source and target domains, thus weakening the domain adaptation per-formance and the decoupling of domain-invariant and domain-specific representations. To solve the aforementioned problems, we propose an image-level domain adaptation framework that aims at adapting source-domain images to the target domain with content-aligned source-target image pairs. Our framework consists of three mutually beneficial modules in a cycle: a cross-domain content alignment module to generate source-target pairs with consistent content representations in a self-supervised manner, a reference-guided image synthesis based on the generated content-aligned source-target image pairs, and a contrastive learning module to self-supervise domain-invariant feature extractor. Our contrastive appearance adaptation is task-agnostic and robust to complex perception tasks in autonomous driving. Our proposed method demonstrates state-of-the-art results in cross-domain object detection, semantic segmentation, and depth estimation as well as better image synthesis ability qualitatively and quantitatively.
Ziqiang Zheng, Yingshu Chen, Binh-Son Hua, Yang Wu 0001, Sai-Kit Yeung
IROS2
2023 CompUDA: Compositional Unsupervised Domain Adaptation for Semantic Segmentation Under Adverse Conditions
abstract
In autonomous driving, performing robust semantic segmentation under adverse weather conditions is a long-standing challenge. Imperfect camera observations under adverse conditions result in images with reduced visibility, which hinders label annotation and semantic scene understanding based on these images. A common solution is to adopt semantic segmentation models trained in a source domain with ground truth labels and perform unsupervised domain adaptation (UDA) from the source domain to an unlabeled target domain that has adverse conditions. Due to imperfect visual observations in the target domain, such adaptation needs special treatment to achieve good performance. In this paper, we propose a new compositional unsupervised domain adaptation (CompUDA) method that disentangles the domain gap based on multiple factors including style, visibility, and image quality. The domain gaps caused by these individual factors can then be addressed separately by introducing the intermediate domains. Specifically, 1) to address the style gap, we perform source-to-intermediate domain adaptation and generate pseudo-labels for self-training in the target domain; 2) to address the visibility gap, we perform a geometry-aligned normal-to-adverse image translation and introduce a synthetic domain; 3) finally, to address the image quality gap between the synthetic and target domain, we perform a synthetic-to-real adaptation based on the generated pseudo-labels. Our compositional unsupervised domain adaptation can be used in conjunction with a wide variety of semantic segmentation methods and result in significant performance improvement across datasets. The codes are available at https://github.com/zhengziqiang/CompUDA.
Ziqiang Zheng, Yingshu Chen, Binh-Son Hua, Sai-Kit Yeung
IROS2
2022 Neural Scene Decoration from a Single Photograph
Hong-Wing Pang, Yingshu Chen, Phuoc-Hieu Le, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
ECCV (23)2
2022 Time-of-Day Neural Style Transfer for Architectural Photographs
abstract
Architectural photography is a genre of photography that focuses on capturing a building or structure in the foreground with dramatic lighting in the background. Inspired by recent successes in image-to-image translation methods, we aim to perform style transfer for architectural photographs. However, the special composition in architectural photography poses great challenges for style transfer in this type of photographs. Existing neural style transfer methods treat the architectural images as a single entity, which would generate mismatched chrominance and destroy geometric features of the original architecture, yielding unrealistic lighting, wrong color rendition, and visual artifacts such as ghosting, appearance distortion, or color mismatching. In this paper, we specialize a neural style transfer method for architectural photography. Our method addresses the composition of the foreground and background in an architectural photograph in a two-branch neural network that separately considers the style transfer of the foreground and the background, respectively. Our method comprises a segmentation module, a learning-based image-to-image translation module, and an image blending optimization module. We trained our image-to-image translation neural network with a new dataset of unconstrained outdoor architectural photographs captured at different magic times of a day, utilizing additional semantic information for better chrominance matching and geometry preservation. Our experiments show that our method can produce photorealistic lighting and color rendition on both the foreground and background, and outperforms general image-to-image translation and arbitrary style transfer baselines quantitatively and qualitatively. Our code and data are available at https://github.com/hkust-vgd/architectural_style_transfer.
Yingshu Chen, Tuan-Anh Vu, Ka-Chun Shum, Sai-Kit Yeung, Binh-Son Hua
ICCP1