Ting-Hsuan Liao

dblp:315/4399 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0001-2295-3502ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 81% Generative modeling · 19%
Computer graphics and multimedia
2 papers
Image and video processing · 77% Visual content generation and editing · 23%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
1.722025
PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos · SIGGRAPH Asia 2025
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework · ICCV 2025
Computer vision › 3D vision
neural radiance field
1.722025
PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos · SIGGRAPH Asia 2025
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework · ICCV 2025
Computer vision › 3D vision › 3d reconstruction
dynamic 3d reconstruction
0.912025
PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos · SIGGRAPH Asia 2025
Computer vision › 3D vision › 3d reconstruction
non-rigid reconstruction
0.912025
PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos · SIGGRAPH Asia 2025
Machine learning › Generative modeling › diffusion model › human motion generation
text-to-motion generation
0.912025
Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions · CVPR 2025
Image and video processing
super-resolution
0.912025
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework · ICCV 2025
Machine learning › Generative modeling
diffusion model
0.312025
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework · ICCV 2025
Machine learning › Generative modeling › cross-modal generation
text-conditioned generation
0.312025
Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions · CVPR 2025
Visual content generation and editing
3d content creation
0.312025
PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos · SIGGRAPH Asia 2025

Methods — techniques the papers use, named apart from their topics

point tracking · 1.7image-to-3d model · 1.7diffusion model · 1.7differentiable rendering · 1.73d gaussian splatting · 1.7variational autoencoder · 0.9language model · 0.9finite scalar quantization · 0.9
YearPublicationVenuePosition
2025 Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions
abstract
We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions. Project URL: https://shape-move.github.io/.
Ting-Hsuan Liao, Yi Zhou 0023, Chun-Hao Paul Huang, Saayan Mitra, Jia-Bin Huang 0001, Uttaran Bhattacharya
CVPR1
2025 Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework
Ting-Hsuan Liao, Pengsheng Guo, Alexander G. Schwing
ICCV2
2025 PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos
abstract
We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences that feature substantial object deformation, large-scale camera movement, and limited view coverage, which typically challenge conventional systems. At its core, our approach trains a personalized, object-centric pose estimator, supervised by a pre-trained image-to-3D model. This guides the optimization of deformable 3D Gaussian representation. The optimization is further regularized by long-term 2D point tracking over the entire input video. By combining generative priors and differentiable rendering, PAD3R reconstructs high-fidelity, articulated 3D representations of objects in a category-agnostic way. Extensive qualitative and quantitative results show that PAD3R is robust and generalizes well across challenging scenarios, highlighting its potential for dynamic scene understanding and 3D content creation. Please refer to our project page for more details: PAD3R.github.io .
Ting-Hsuan Liao, Songwei Ge, Gengshan Yang, Jia-Bin Huang 0001
SIGGRAPH Asia1
2023 Pixel-Wise Prediction based Visual Odometry via Uncertainty Estimation
abstract
This paper introduces pixel-wise prediction based visual odometry (PWVO), which is a dense prediction task that evaluates the values of translation and rotation for every pixel in its input observations. PWVO employs uncertainty estimation to identify the noisy regions in the input observations, and adopts a selection mechanism to integrate pixel-wise predictions based on the estimated uncertainty maps to derive the final translation and rotation. In order to train PWVO in a comprehensive fashion, we further develop a data generation workflow for generating synthetic training data. The experimental results show that PWVO is able to deliver favorable results. In addition, our analyses validate the effectiveness of the designs adopted in PWVO, and demonstrate that the uncertainty maps estimated by PWVO is capable of capturing the noises in its input observations.
Hao-Wei Chen, Ting-Hsuan Liao, Hsuan-Kung Yang, Chun-Yi Lee
WACV2
2022 ELDA: Using Edges to Have an Edge on Semantic Segmentation Based UDA
Ting-Hsuan Liao, Huang-Ru Liao, Shan-Ya Yang, Jie-En Yao, Li-Yuan Tsao, Hsu-Shen Liu, Chen-Hao Chao, Bo-Wun Cheng, Chia-Che Chang, Yi-Chen Lo, Chun-Yi Lee
BMVC1
2022 Investigation of Factorized Optical Flows as Mid-Level Representations
abstract
In this paper, we introduce a new concept of incorporating factorized flow maps as mid-level representations, for bridging the perception and the control modules in modular learning based robotic frameworks. To investigate the advantages of factorized flow maps and examine their interplay with the other types of mid-level representations, we further develop a configurable framework, along with four different environments that contain both static and dynamic objects, for analyzing the impacts of factorized optical flow maps on the performance of deep reinforcement learning agents. Based on this framework, we report our experimental results on various scenarios, and offer a set of analyses to justify our hypothesis. Finally, we validate flow factorization in real world scenarios.
Hsuan-Kung Yang, Tsu-Ching Hsiao, Ting-Hsuan Liao, Hsu-Shen Liu, Li-Yuan Tsao, Tzu-Wen Wang, Shan-Ya Yang, Huang-Ru Liao, Chun-Yi Lee
IROS3