Xinyao Liao

dblp:365/4407 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 One-Stage Absolute Human Mesh Recovery
abstract
The reconstruction of realistic and precise human meshes in world coordinates is facilitated by considering scene information. Challenges related to accuracy, robustness, and computation time are faced by existing absolute human mesh recovery methods. In this paper, a one-stage model for absolute human mesh recovery with superior reconstruction precision and inference speed is presented. The proposed one-stage model is composed of two parallel branches to achieve root position estimation and human mesh regression. To effectively connect the two branches, a scene-image information aggregation module is designed. The accuracy of the estimated human meshes is improved and the end-to-end training of the whole model is facilitated by this module. Experiments are conducted on three diverse datasets, and a GMPJPE decrease of 72.3 mm/27.32% and an MPJPE reduction of 25.6 mm/27.26% are achieved by the proposed method with the lowest inference time compared to previous SOTA methods.
Xinyao Liao, Wanjuan Su, Chen Zhang 0043, Ximeng Li 0007, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.1
2025 High-Fidelity Lightweight Mesh Reconstruction from Point Clouds
abstract
Recently, learning signed distance functions (SDFs) from point clouds has become popular for reconstruction. To ensure accuracy, most methods require using high-resolution Marching Cubes for surface extraction. However, this results in redundant mesh elements, making the mesh inconvenient to use. To solve the problem, we propose an adaptive meshing method to extract resolution-adaptive meshes based on surface curvature, enabling the recovery of high-fidelity lightweight meshes. Specifically, we first use point-based representation to perceive implicit surfaces and calculate surface curvature. A vertex generator is designed to produce curvature-adaptive vertices with any specified number on the implicit surface, preserving the overall structure and high-curvature features. Then we develop a Delaunay meshing algorithm to generate meshes from vertices, ensuring geometric fidelity and correct topology. In addition, to obtain accurate SDFs for adaptive meshing and achieve better lightweight reconstruction, we design a hybrid representation combining feature grid and feature tri-plane for better detail capture. Experiments demonstrate that our method can generate high-quality lightweight meshes from point clouds. Compared with methods from various categories, our approach achieves superior results, especially in capturing more details with fewer elements.
Chen Zhang 0043, Ximeng Li 0007, Xinyao Liao, Wanjuan Su, Wenbing Tao
CVPR4
2025 Motionagent: Fine-Grained Controllable Video Generation via Motion Field Agent
abstract
We propose MotionAgent, enabling fine-grained motion control for text-guided image-to-video generation. The key technique is the motion field agent that converts motion information in text prompts into explicit motion fields, providing flexible and precise motion guidance. Specifically, the agent extracts the object movement and camera motion described in the text and converts them into object trajectories and camera extrinsics, respectively. An analytical optical flow composition module integrates these motion representations in 3D space and projects them into a unified optical flow. An optical flow adapter takes the flow to control the base image-to-video diffusion model for generating fine-grained controlled videos. The significant improvement in the Video-Text Camera Motion metrics on VBench indicates that our method achieves precise control over camera motion. We construct a subset of VBench to evaluate the alignment of motion information in the text and the generated video, outperforming other advanced models on motion generation accuracy.
Xinyao Liao, Xianfang Zeng, Gang Yu 0002, Guosheng Lin, Chi Zhang 0007
ICCV1
2025 InstaHMR: Instance-Aware One-Stage Multi-Person Human Mesh Recovery
abstract
Human mesh recovery aims to estimate all human meshes within a given image. In this article, we propose an Instance-aware Multi-person 3D Human Mesh Recovery (InstaHMR) network based on the one-stage framework. Compared to former one-stage methods, instance-aware single person feature is exploited to represent more accurate human mesh. Specifically, we propose the Contextual Instance Guidance (CIG) module which generates instance-aware single person feature by leveraging spatial and channel attention operations. In this way, it preserves more instance-specific information compared to the pixel-level feature used in some existing one-stage methods. Besides, we further introduce two auxiliary losses for better mesh recovery, namely the Human Triplet Planes (HTP) loss and the T-pose Shape (TS) loss. The HTP loss encourages the model to capture subtle differences in human joint positions, while the TS loss facilitates the learning of abstract shape parameters. By incorporating these advancements, our model achieves state-of-the-art results on four multi-person datasets.
Xinyao Liao, Chen Zhang 0043, Jianyao Xu, Wanjuan Su, Zhi Chen 0011, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.1
2025 PG-NeuS: Robust and Efficient Point Guidance for Multi-View Neural Surface Reconstruction
abstract
Recently, learning multi-view neural surface reconstruction with the supervision of point clouds or depth maps has been a promising way. However, due to weak perception and underutilization of prior information, current methods still struggle with the challenges of limited accuracy and excessive time complexity. In addition, prior data perturbation is also an important yet rarely considered issue, often resulting in distorted geometry. To address these challenges, we propose a novel point-guided method named PG-NeuS, which achieves accurate and efficient reconstruction while robustly coping with point noise. Specifically, the aleatoric uncertainty of the point cloud is modeled to capture the noise distribution, estimating the reliability of each point and enhancing robustness against noise. Moreover, a Neural Projection module is proposed to connect points and images, adding geometric constraints to the implicit surface and achieving more precise point guidance. To better compensate for geometric bias between volume rendering and point modeling, we additionally design a Bias network that leverages the geometric information in high-fidelity points to enhance detail representation. Benefiting from the effective point guidance, the proposed PG-NeuS achieves an 11x speed increase and a 33.3% accuracy improvement compared to NeuS on DTU, even with a lightweight network. Extensive experiments show that our method yields high-quality surfaces with high efficiency, especially for fine-grained details and smooth regions, outperforming the state-of-the-art methods. Moreover, it exhibits strong robustness to noisy data and sparse data.
Chen Zhang 0043, Wanjuan Su, Qingshan Xu 0001, Xinyao Liao, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.4
2024 UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
abstract
Scene Graph Generation(SGG) is a scene understanding task that aims at identifying object entities and reasoning their relationships within a given image. In contrast to prevailing two-stage methods based on a large object detector (e.g., Faster R-CNN), one-stage methods integrate a fixed-size set of learnable queries to jointly reason relational triplets . This paradigm demonstrates robust performance with significantly reduced parameters and computational overhead. However, the challenge in one-stage methods stems from the issue of weak entanglement, wherein entities involved in relationships require both coupled features shared within triplets and decoupled visual features. Previous methods either adopt a single decoder for coupled triplet feature modeling or multiple decoders for separate visual feature extraction but fail to consider both. In this paper, we introduce UniQ, a Unified decoder with task-specific Queries architecture, where task-specific queries generate decoupled visual features for subjects, objects, and predicates respectively, and unified decoder enables coupled feature modeling within relational triplets. Experimental results on the Visual Genome dataset demonstrate that UniQ has superior performance to both one-stage and two-stage methods.
Xinyao Liao, Wei Wei 0002, Dangyang Chen
ACM Multimedia1