Sheng-Yu Huang

dblp:82/10068 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-3149-9620ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 3D Gaussian Inpainting with Depth-Guided Cross-View Consistency
abstract
When performing 3D inpainting using novel-view rendering methods like Neural Radiance Field (NeRF) or 3D Gaussian Splatting (3DGS), how to achieve texture and geometry consistency across camera views has been a challenge. In this paper, we propose a framework of 3D Gaussian Inpainting with Depth-Guided Cross-View Consistency (3DGIC) for cross-view consistent 3D inpainting. Guided by the rendered depth information from each training view, our 3DGIC exploits background pixels visible across different views for updating the inpainting mask, allowing us to refine the 3DGS for inpainting purposes. Through extensive experiments on benchmark datasets, we confirm that our 3DGIC outperforms current state-of-the-art 3D inpainting methods quantitatively and qualitatively.
Sheng-Yu Huang, Zi-Ting Chou, Yu-Chiang Frank Wang
CVPR1
2025 Data-Efficient 3D Visual Grounding via Order-Aware Referring
abstract
3D visual grounding aims to identify the target object within a 3D point cloud scene referred to by a natural language description. Previous works usually require significant data relating to point color and their descriptions to exploit the corresponding complicated verbo-visual relations. In our work, we introduce Vigor, a novel Data-Efficient 3D Visual Grounding framework via Order-aware Referring. Vigor leverages LLM to produce a desirable referential order from the input description for 3D visual grounding. With the proposed stacked object-referring blocks, the predicted anchor objects in the above order allow one to locate the target object progressively with-out supervision on the identities of anchor objects or exact relations between anchor/target objects. We also present an order-aware warm-up training strategy, which augments referential orders for pre-training the visual grounding framework, allowing us to better capture the complex verbo-visual relations and benefit the desirable data-efficient learning scheme. Experimental results on the NR3D and ScanRefer datasets demonstrate our superiority in low-resource scenarios. In particular, Vigor surpasses current state-of-the-art frameworks by 9.3% and 7.6% grounding accuracy under 1% data and 10% data settings on the NR3D dataset, respectively.
Tung-Yu Wu, Sheng-Yu Huang, Yu-Chiang Frank Wang
WACV2
2025 Learning Shape-Color Diffusion Priors for Text-Guided 3D Object Generation
abstract
Generating 3D shapes according to specific textual input is a crucial topic in the multimedia application, with its potential enhancement to the VR/AR/XR usage that enables more diverse virtual scenes. Due to the recent success of diffusion models, text-guided 3D object generation has drawn a lot of attention recently. However, current latent diffusion-based methods are restricted to shape-only generation, requiring time-consuming and computationally expensive post-processing to obtain colored objects. In this paper, we propose an end-to-endShape-Color Diffusion Prior framework (SCDiff)to achieve colored text-to-3D object generation. Given a general text description as input, our SCDiff is able to distinguish shape and color-related priors in the text and generate a shape latent and a color latent for a pre-trained 3D object auto-encoder to derive colored 3D objects. Our SCDiff contains two 3D latent diffusion models (LDM), where one generates the shape latent from the input text and the other generates the color latent. To help the two LDMs focus on shape/color-related information, we further adopt a Large Language Model (LLM) to separate the input text into a shape phrase and a color phrase via an in-context learning technique so that our shape/color LDM would not be influenced by irrelevant information. Due to the separation of shape and color latent, we are able to manipulate the color of an object by giving different color phrases while maintaining the original shape. Experiments on a benchmark dataset would quantitatively and qualitatively verify the effectiveness and practicality of our proposed model. As an extension, we show the capability of our SCDiff on 3D object generation and manipulation based on various modality conditions, which further confirms the scalability and applications in multimedia of our proposed framework.
Sheng-Yu Huang, Chi-Pin Huang, Kai-Po Chang, Zi-Ting Chou, I-Jieh Liu, Yu-Chiang Frank Wang
IEEE Trans. Multim.1
2024 GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene Understanding
abstract
Utilizing multi-view inputs to synthesize novel-view images, Neural Radiance Fields (NeRF) have emerged as a popular research topic in 3D vision. In this work, we introduce a Generalizable Semantic Neural Radiance Fields (GSNeRF), which uniquely takes image semantics into the synthesis process so that both novel view image and the associated semantic maps can be produced for unseen scenes. Our GSNeRF is composed of two stages: Semantic Geo-Reasoning and Depth-Guided Visual rendering. The former is able to observe multi-view image inputs to extract semantic and geometry features from a scene. Guided by the resulting image geometry information, the latter performs both image and semantic rendering with improved performances. Our experiments not only confirm that GSNeRF performs favorably against prior works on both novel-view image and semantic segmentation synthesis but the effectiveness of our sampling strategy for visual rendering is further verified.
Zi-Ting Chou, Sheng-Yu Huang, I-Jieh Liu, Yu-Chiang Frank Wang
CVPR2
2024 TPA3D: Triplane Attention for Fast Text-to-3D Generation
Bin-Shih Wu, Hong-En Chen, Sheng-Yu Huang, Yu-Chiang Frank Wang
ECCV (18)3
2023 Interpreting Latent Representation in Neural Radiance Fields for Manipulating Object Semantics
abstract
Manipulating 3D objects has been among the active research topic for 3D vision. With the development and success of neural radiance field (NeRF) [1] on scene modeling, synthesizing and manipulating 3D objects using such a representation becomes desirable. In this paper, we introduce a semantic-aware generative NeRF, which is able to interpret the latent representation learned by category-specific generative NeRFs and to achieve editing of particular part attributes. With pretrained generative NeRF, we propose to deploy a semantic segmentor for performing part segmentation on the object category. This allows the rendering of the 2D image and prediction of the corresponding segmentation mask. Our proposed scheme learns to manipulate the resulting latent representation, optimized to edit the object part of interest with varying degrees. We conduct experiments on various object categories on benchmark datasets, and the results successfully verify the effectiveness and practicality of our proposed model.
Yu-Shan Huang, Sheng-Yu Huang, Hao-Yu Hsu, Yu-Chiang Frank Wang
ICIP2
2022 3D-Selfcutmix: Self-Supervised Learning for 3D Point Cloud Analysis
abstract
Point clouds have been widely applied to represent 3D data, with a variety of applications such as autonomous driving, augmented reality, and robotics. Since collecting a large amount of labeled 3D point cloud data for training deep learning models might not always be applicable, we propose the novel learning strategy of 3D-SelfCutMix, which advances mixed-sample data augmentation techniques while exploiting the spatial and semantic consistencies between point cloud data. Depending on the availability of label supervision, the proposed network can be realized in either self-supervised or fully-supervised manners, while both versions are shown to benefit downstream tasks. In our experiments, we consider a variety of tasks including classification and part-segmentation tasks, which sufficiently support the use of the proposed method for 3D point cloud analysis.
Yuan-Yi Xu, Yan-Yang Ji, Sheng-Yu Huang, Zhi-Hao Lin, Yu-Chiang Frank Wang
ICIP3
2022 SPoVT: Semantic-Prototype Variational Transformer for Dense Point Cloud Semantic Completion
abstract
Point cloud completion is an active research topic for 3D vision and has been widelystudied in recent years. Instead of directly predicting missing point cloud fromthe partial input, we introduce a Semantic-Prototype Variational Transformer(SPoVT) in this work, which takes both partial point cloud and their semanticlabels as the inputs for semantic point cloud object completion. By observingand attending at geometry and semantic information as input features, our SPoVTwould derive point cloud features and their semantic prototypes for completionpurposes. As a result, our SPoVT not only performs point cloud completion withvarying resolution, it also allows manipulation of different semantic parts of anobject. Experiments on benchmark datasets would quantitatively and qualitativelyverify the effectiveness and practicality of our proposed model.
Sheng-Yu Huang, Hao-Yu Hsu, Yu-Chiang Frank Wang
NeurIPS1
2022 Learning of 3D Graph Convolution Networks for Point Cloud Analysis
abstract
Point clouds are among the popular geometry representations in 3D vision. However, unlike 2D images with pixel-wise layouts, such representations containing unordered data points which make the processing and understanding the associated semantic information quite challenging. Although a number of previous works attempt to analyze point clouds and achieve promising performances, their performances would degrade significantly when data variations like shift and scale changes are presented. In this paper, we propose 3D graph convolution networks (3D-GCN), which uniquely learns 3D kernels with graph max-pooling mechanisms for extracting geometric features from point cloud data across different scales. We show that, with the proposed 3D-GCN, satisfactory shift and scale invariance can be jointly achieved. We show that 3D-GCN can be applied to point cloud classification and segmentation tasks, with ablation studies and visualizations verifying the design of 3D-GCN.
Zhi-Hao Lin, Sheng-Yu Huang, Yu-Chiang Frank Wang
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Convolution in the Cloud: Learning Deformable Kernels in 3D Graph Convolution Networks for Point Cloud Analysis
abstract
Point clouds are among the popular geometry representations for 3D vision applications. However, without regular structures like 2D images, processing and summarizing information over these unordered data points are very challenging. Although a number of previous works attempt to analyze point clouds and achieve promising performances, their performances would degrade significantly when data variations like shift and scale changes are presented. In this paper, we propose 3D Graph Convolution Networks (3D-GCN), which is designed to extract local 3D features from point clouds across scales, while shift and scale-invariance properties are introduced. The novelty of our 3D-GCN lies in the definition of learnable kernels with a graph max-pooling mechanism. We show that 3D-GCN can be applied to 3D classification and segmentation tasks, with ablation studies and visualizations verifying the design of 3D-GCN.
Zhi-Hao Lin, Sheng-Yu Huang, Yu-Chiang Frank Wang
CVPR2
2010 Emotion recognition based on a novel triangular facial feature extraction method
abstract
Recognizing human emotions from facial expressions is highly dependent on the quality of the referred facial expression features. Conventional methods often suffer from high computation time and serious influence of environment variations. In this paper, a triangular facial feature extraction method based on a Modified Active Shape Model (MASM) is proposed. This method features considering the interactions of all facial features, escaping from the affection of environment variations as well as noisy facial features, and reducing feature dimensions. MASM adopts the same shape representation and shape training procedures as ASM, but executes a different landmark searching procedure without using the gray level training procedure to avoid the affection from environment variations. Using the feature points extracted by MASM, two methods, one is based on statistical analysis and another one is derived from the genetic algorithm, are proposed to extract an optimal set of triangular facial features for emotion recognition. In the experiments with JAFFE database, a neural network classifier is employed to recognize emotions with those extracted triangular facial features. The experimental results show that based on the statistical analysis 65.1% recognition rate is achieved, and based on the genetic algorithm 70.2% recognition rate is achieved.
Kuan-Chieh Huang, Sheng-Yu Huang, Yau-Hwang Kuo
IJCNN2