Shih-Yang Su

dblp:201/7133 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 69% Face, body and person analysis · 10% Reinforcement learning · 9%
Computer graphics and multimedia
4 papers
Geometric modeling and processing · 52% Rendering · 41% Image and video processing · 7%

Topics — the 26 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
neural radiance field
1.422025
Locality Sensitive Avatars From Video · ICLR 2025
A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose · NeurIPS 2021
Computer vision › 3D vision › 3d shape representation
canonical representation
0.912025
Locality Sensitive Avatars From Video · ICLR 2025
Computer vision › 3D vision › 3d human reconstruction
human avatar reconstruction
0.912025
Locality Sensitive Avatars From Video · ICLR 2025
Geometric modeling and processing
deformation modeling
0.912025
Locality Sensitive Avatars From Video · ICLR 2025
Geometric modeling and processing › shape deformation
non-rigid deformation
0.912025
Locality Sensitive Avatars From Video · ICLR 2025
Rendering
neural rendering
0.812024
Gaussian Shadow Casting for Neural Characters · CVPR 2024
Rendering
relighting
0.812024
Gaussian Shadow Casting for Neural Characters · CVPR 2024
Computer vision › 3D vision › human body modeling › 3d human modeling
animatable human avatar
0.712023
NPC: Neural Point Characters from Video · ICCV 2023
Machine learning › Reinforcement learning
deep reinforcement learning
0.722018
Diversity-Driven Exploration Strategy for Deep Reinforcement Learning · NeurIPS 2018
Virtual-to-Real: Learning to Control in Visual Semantic Segmentation · IJCAI 2018
Computer vision › 3D vision › motion capture
human performance capture
0.712023
NPC: Neural Point Characters from Video · ICCV 2023
Computer vision › 3D vision
3d human reconstruction
0.612022
DANBO: Disentangled Articulated Neural Body Representations via Graph Neural Networks · ECCV (2) 2022
Computer vision › Face, body and person analysis › human pose estimation
articulated body model
0.612022
DANBO: Disentangled Articulated Neural Body Representations via Graph Neural Networks · ECCV (2) 2022
Machine learning › Graph learning
graph neural network
0.612022
DANBO: Disentangled Articulated Neural Body Representations via Graph Neural Networks · ECCV (2) 2022
Computer vision › Face, body and person analysis
human pose estimation
0.512021
A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose · NeurIPS 2021
Computer vision › 3D vision
motion capture
0.512021
A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose · NeurIPS 2021
Computer vision › 3D vision › depth estimation
layered depth image
0.412020
3D Photography Using Context-Aware Layered Depth Inpainting · CVPR 2020
Computer vision › 3D vision
novel view synthesis
0.412020
3D Photography Using Context-Aware Layered Depth Inpainting · CVPR 2020
Machine learning › Reinforcement learning › exploration
exploration strategies
0.312018
Diversity-Driven Exploration Strategy for Deep Reinforcement Learning · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.312018
Virtual-to-Real: Learning to Control in Visual Semantic Segmentation · IJCAI 2018
Robotics › Robot navigation and mapping
visual navigation
0.312018
Virtual-to-Real: Learning to Control in Visual Semantic Segmentation · IJCAI 2018
Computer vision › 3D vision
3d reconstruction
0.212023
NPC: Neural Point Characters from Video · ICCV 2023
Computer vision › 3D vision
canonicalization
0.212023
NPC: Neural Point Characters from Video · ICCV 2023
Geometric modeling and processing › shape modeling
human body model
0.112021
A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose · NeurIPS 2021
Image and video processing › image restoration › image inpainting
depth-aware inpainting
0.112020
3D Photography Using Context-Aware Layered Depth Inpainting · CVPR 2020
Image and video processing › image restoration
image inpainting
0.112020
3D Photography Using Context-Aware Layered Depth Inpainting · CVPR 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112018
Virtual-to-Real: Learning to Control in Visual Semantic Segmentation · IJCAI 2018

Methods — techniques the papers use, named apart from their topics

graph neural network · 2.3neural radiance field · 1.7gaussian splatting · 1.5gaussian density proxy · 1.5deferred neural rendering · 1.5inverse kinematics · 1.0template-free reconstruction · 0.7physics-free deformation · 0.7neural point representation · 0.7disentangled representation · 0.6reparameterization · 0.5learning-based inpainting · 0.4layered depth inpainting · 0.4
YearPublicationVenuePosition
2025 Locality Sensitive Avatars From Video
abstract
We present locality-sensitive avatar, a neural radiance field (NeRF) based network to learn human motions from monocular videos. To this end, we estimate a canonical representation between different frames of a video with a non-linear mapping from observation to canonical space, which we decompose into a skeletal rigid motion and a non-rigid counterpart. Our key contribution is to retain fine-grained details by modeling the non-rigid part with a graph neural network (GNN) that keeps the pose information local to neighboring body parts. Compared to former canonical representation based methods which solely operate on the coordinate space of a whole shape, our locality-sensitive motion modeling can reproduce both realistic shape contours and vivid fine-grained details. We evaluate on ZJU-MoCap, SynWild, ActorsHQ, MVHumanNet and various outdoor videos. The experiments reveal that with the locality sensitive deformation to canonical feature space, we are the first to achieve state-of-the-art results across novel view synthesis, novel pose animation and 3D shape reconstruction simultaneously. Our code is available at https://github.com/ChunjinSong/lsavatar.
Chunjin Song, Shih-Yang Su, Bastian Wandt, Leonid Sigal, Helge Rhodin
ICLR3
2024 Mirror-Aware Neural Humans
abstract
Human motion capture either requires multi-camera systems or is unreliable when using single-view input due to depth ambiguities. Meanwhile, mirrors are readily available in urban environments and form an affordable alternative by recording two views with only a single camera. However, the mirror setting poses the additional challenge of handling occlusions of real and mirror image. Going beyond existing mirror approaches for 3D human pose estimation, we utilize mirrors for learning a complete body model, including shape and dense appearance. Our main contributions are extending articulated neural radiance fields to include a notion of a mirror, making it sample-efficient over potential occlusion regions. Together, our contributions realize a consumer-level 3D motion capture system that starts from off-the-shelf 2D poses by automatically calibrating the camera, estimating mirror orientation, and subsequently lifting 2D keypoint detections to 3D skeleton pose that is used to condition the mirror-aware NeRF. We empirically demonstrate the benefit of learning a body model and accounting for occlusion in challenging mirror scenes. The project is available at: https://danielajisafe.github.io/mirror-aware-neural-humans/.
Daniel Ajisafe, James Tang, Shih-Yang Su, Bastian Wandt, Helge Rhodin
3DV3
2024 Gaussian Shadow Casting for Neural Characters
abstract
Neural character models can now reconstruct detailed geometry and texture from video, but they lack explicit shadows and shading, leading to artifacts when generating novel views and poses or during relighting. It is particularly difficult to include shadows as they are a global effect and the required casting of secondary rays is costly. We propose a new shadow model using a Gaussian density proxy that replaces sampling with a simple analytic formula. It supports dynamic motion and is tailored for shadow computation, thereby avoiding the affine projection approximation and sorting required by the closely related Gaussian splatting. Combined with a deferred neural rendering model, our Gaussian shadows enable Lambertian shading and shadow casting with minimal overhead. We demonstrate improved reconstructions, with better separation of albedo, shading, and shadows in challenging outdoor scenes with direct sun light and hard shadows. Our method is able to optimize the light direction without any input from the user. As a result, novel poses have fewer shadow artifacts, and relighting in novel scenes is more realistic compared to the state-of-the-art methods, providing new ways to pose neural characters in novel environments, increasing their applicability. Code available at: https://github.com/LuisBolanos17/GaussianShadowCasting
Luis Bolanos, Shih-Yang Su, Helge Rhodin
CVPR2
2023 NPC: Neural Point Characters from Video
abstract
High-fidelity human 3D models can now be learned directly from videos, typically by combining a template-based surface model with neural representations. However, obtaining a template surface requires expensive multi-view capture systems, laser scans, or strictly controlled conditions. Previous methods avoid using a template but rely on a costly or ill-posed mapping from observation to canonical space. We propose a hybrid point-based representation for reconstructing animatable characters that does not require an explicit surface model, while being generalizable to novel poses. For a given video, our method automatically produces an explicit set of 3D points representing approximate canonical geometry, and learns an articulated deformation model that produces pose-dependent point transformations. The points serve both as a scaffold for high-frequency neural features and an anchor for efficiently mapping between observation and canonical space. We demonstrate on established benchmarks that our representation overcomes limitations of prior work operating in either canonical or in observation space. Moreover, our automatic point extraction approach enables learning models of human and animal characters alike, matching the performance of the methods using rigged surface templates despite being more general. Project website: https://lemonatsu.github.io/npc/.
Shih-Yang Su, Timur M. Bagautdinov, Helge Rhodin
ICCV1
2022 DANBO: Disentangled Articulated Neural Body Representations via Graph Neural Networks
Shih-Yang Su, Timur M. Bagautdinov, Helge Rhodin
ECCV (2)1
2021 A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose
abstract
While deep learning reshaped the classical motion capture pipeline with feed-forward networks, generative models are required to recover fine alignment via iterative refinement. Unfortunately, the existing models are usually hand-crafted or learned in controlled conditions, only applicable to limited domains. We propose a method to learn a generative neural body model from unlabelled monocular videos by extending Neural Radiance Fields (NeRFs). We equip them with a skeleton to apply to time-varying and articulated motion. A key insight is that implicit models require the inverse of the forward kinematics used in explicit surface models. Our reparameterization defines spatial latent variables relative to the pose of body parts and thereby overcomes ill-posed inverse operations with an overparameterization. This enables learning volumetric body shape and appearance from scratch while jointly refining the articulated pose; all without ground truth labels for appearance, pose, or 3D shape on the input videos. When used for novel-view-synthesis and motion capture, our neural model improves accuracy on diverse datasets.
Shih-Yang Su, Frank Yu, Michael Zollhöfer, Helge Rhodin
NeurIPS1
2020 3D Photography Using Context-Aware Layered Depth Inpainting
abstract
We propose a method for converting a single RGB-D input image into a 3D photo, i.e., a multi-layer representation for novel view synthesis that contains hallucinated color and depth structures in regions occluded in the original view. We use a Layered Depth Image with explicit pixel connectivity as underlying representation, and present a learning-based inpainting model that iteratively synthesizes new local color-and-depth content into the occluded region in a spatial context-aware manner. The resulting 3D photos can be efficiently rendered with motion parallax using standard graphics engines. We validate the effectiveness of our method on a wide range of challenging everyday scenes and show less artifacts when compared with the state-of-the-arts.
Meng-Li Shih, Shih-Yang Su, Johannes Kopf 0001, Jia-Bin Huang 0001
CVPR2
2018 Virtual-to-Real: Learning to Control in Visual Semantic Segmentation
abstract
Collecting training data from the physical world is usually time-consuming and even dangerous for fragile robots, and thus, recent advances in robot learning advocate the use of simulators as the training platform. Unfortunately, the reality gap between synthetic and real visual data prohibits direct migration of the models trained in virtual worlds to the real world. This paper proposes a modular architecture for tackling the virtual-to-real problem. The proposed architecture separates the learning model into a perception module and a control policy module, and uses semantic image segmentation as the meta representation for relating these two modules. The perception module translates the perceived RGB image to semantic image segmentation. The control policy module is implemented as a deep reinforcement learning agent, which performs actions based on the translated image segmentation. Our architecture is evaluated in an obstacle avoidance task and a target following task. Experimental results show that our architecture significantly outperforms all of the baseline methods in both virtual and real environments, and demonstrates a faster learning curve than them. We also present a detailed analysis for a variety of variant configurations, and validate the transferability of our modular architecture.
Zhang-Wei Hong, Yu-Ming Chen 0002, Hsuan-Kung Yang, Shih-Yang Su, Tzu-Yun Shann, Yi-Hsiang Chang, Brian Hsi-Lin Ho, Chih-Chieh Tu, Tsu-Ching Hsiao, Hsin-Wei Hsiao, Sih-Pin Lai, Yueh-Chuan Chang, Chun-Yi Lee
IJCAI4
2018 Diversity-Driven Exploration Strategy for Deep Reinforcement Learning
abstract
Efficient exploration remains a challenging research problem in reinforcement learning, especially when an environment contains large state spaces, deceptive local optima, or sparse rewards. To tackle this problem, we present a diversity-driven approach for exploration, which can be easily combined with both off- and on-policy reinforcement learning algorithms. We show that by simply adding a distance measure to the loss function, the proposed methodology significantly enhances an agent's exploratory behaviors, and thus preventing the policy from being trapped in local optima. We further propose an adaptive scaling method for stabilizing the learning process. We demonstrate the effectiveness of our method in huge 2D gridworlds and a variety of benchmark environments, including Atari 2600 and MuJoCo. Experimental results show that our method outperforms baseline approaches in most tasks in terms of mean scores and exploration efficiency.
Zhang-Wei Hong, Tzu-Yun Shann, Shih-Yang Su, Yi-Hsiang Chang, Tsu-Jui Fu, Chun-Yi Lee
NeurIPS3
2017 Automatic conversion of Pop music into chiptunes for 8-bit pixel art
abstract
In this paper, we propose an audio mosaicing method that converts Pop songs into a specific music style called “chiptune,” or “8-bit music.” The goal is to reproduce Pop songs by using the sound of the chips on the old game consoles in 1980s/1990s. The proposed method goes through a procedure that first analyzes the pitches of an incoming Pop song in the frequency domain, and then synthesizes the song with template waveforms in the time domain to make it sound like 8-bit music. Because a Pop song is usually composed of the vocal melody and the instrumental accompaniment, in the analysis stage we use a singing voice separation algorithm to separate the vocals from the instruments, and then apply different pitch detection algorithms to transcribe the two separated sources. We validate through a subjective listening test that the proposed method creates much better 8-bit music than existing nonnegative matrix factorization based methods can do. Moreover, we find that synthesis in the time domain is important for this task.
Shih-Yang Su, Cheng-Kai Chiu, Li Su 0004, Yi-Hsuan Yang
ICASSP1