Tianyi Xie

dblp:161/7169 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PhysMotion: Physics-Grounded Dynamics From a Single Image
abstract
We introduce PhysMotion, a novel framework that leverages principled physics-based simulations to guide intermediate 3D representations generated from a single image and input conditions (e.g., applied force and torque), producing high-quality, physically plausible video generation. By utilizing continuum mechanics-based simulations as a prior knowledge, our approach addresses the limitations of traditional data-driven generative models and results in more consistent physically plausible motions. Our framework begins by reconstructing a feed-forward 3D Gaussian from a single image through geometry optimization. This representation is then time-stepped using Material Point Method (MPM) with continuum mechanics-based elastoplasticity models, which provides a strong foundation for realistic$d y$namics, albeit at a coarse level of detail. To enhance the geometry, appearance, and ensure spatiotemporal consistency, we refine the initial simulation using a text-to-image (T2I) diffusion model with cross-frame attention, resulting in a physically plausible video that retains intricate details comparable to the input image. We conduct comprehensive qualitative and quantitative evaluations to validate the efficacy of our method.
Xiyang Tan, Xuan Li 0015, Zeshun Zong, Tianyi Xie, Yin Yang 0002, Chenfanfu Jiang
3DV5
2025 GarmentDreamer: 3DGS Guided Garment Synthesis with Diverse Geometry and Texture Details
abstract
Traditional 3D garment creation is labor-intensive, involving sketching, modeling, UV mapping, and texturing, which are time-consuming and costly. Recent advances in diffusion-based generative models have enabled new possibilities for 3D garment generation from text prompts, images, and videos. However, existing methods either suffer from inconsistencies among multi-view images or require additional processes to separate cloth from the underlying human model. In this paper, we propose GarmentDreamer, a novel method that leverages 3D Gaussian Splatting (GS) as guidance to generate wearable, simulation-ready 3D garment meshes from text prompts. In contrast to using multi-view images directly predicted by generative models as guidance, our 3DGS guidance ensures consistent optimization in both garment deformation and texture synthesis. Our method introduces a novel garment augmentation module, guided by normal and RGBA information, and employs implicit Neural Texture Fields (NeTF) combined with Variational Score Distillation (VSD) to generate diverse geometric and texture details. We validate the effectiveness of our approach through comprehensive qualitative and quantitative experiments, showcasing the superior performance of GarmentDreamer over state-of-the-art alternatives11Demos and codes are available at https://xuan-li.github.io/GarmentDreamerDemo/.
Boqian Li, Xuan Li 0015, Tianyi Xie, Feng Gao 0013, Huamin Wang 0001, Yin Yang 0002, Chenfanfu Jiang
3DV4
2025 PhysAnimator: Physics-Guided Generative Cartoon Animation
abstract
Creating hand-drawn animation sequences is laborintensive and demands professional expertise. We introduce PhysAnimator, a novel approach for generating physically plausible meanwhile anime-stylized animation from static anime illustrations. Our method seamlessly integrates physics-based simulations with data-driven generative models to produce dynamic and visually compelling animations. To capture the fluidity and exaggeration characteristic of anime, we perform image-space deformable body simulations on extracted mesh geometries. We enhance artistic control by introducing customizable energy strokes and incorporating rigging point support, enabling the creation of tailored animation effects such as wind interactions. Finally, we extract and warp sketches from the simulation sequence, generating a texture-agnostic representation, and employ a sketch-guided video diffusion model to synthesize high-quality animation frames. The resulting animations exhibit temporal consistency and visual plausibility, demonstrating the effectiveness of our method in creating dynamic anime-style animations. See our project page for more demos: https://xpandora.github.io/PhysAnimator/.
Tianyi Xie, Chenfanfu Jiang
CVPR1
2025 VideoPhy: Evaluating Physical Commonsense for Video Generation
abstract
Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of visual concepts, synthesize realistic motions and render complex objects. Hence, these generative models have the potential to become general-purpose simulators of the physical world. However, it is unclear how far we are from this goal with the existing text-to-video generative models. To this end, we present VideoPhy, a benchmark designed to assess whether the generated videos follow physical commonsense for real-world activities (e.g. marbles will roll down when placed on a slanted surface). Specifically, we curate diverse prompts that involve interactions between various material types in the physical world (e.g., solid-solid, solid-fluid, fluid-fluid). We then generate videos conditioned on these captions from diverse state-of-the-art text-to-video generative models, including open models (e.g., CogVideoX) and closed models (e.g., Lumiere, Dream Machine). Our human evaluation reveals that the existing models severely lack the ability to generate videos adhering to the given text prompts, while also lack physical commonsense. Specifically, the best performing model, CogVideoX-5B, generates videos that adhere to the caption and physical laws for 39.6% of the instances. VideoPhy thus highlights that the video generative models are far from accurately simulating the physical world. Finally, we propose an auto-evaluator, VideoCon-Physics, to assess the performance reliably for the newly released models. The code is available here: https://github.com/Hritikbansal/videophy.
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang 0001, Aditya Grover
ICLR3
2025 GRIP: A General Robotic Incremental Potential Contact Simulation Dataset for Unified Deformable-Rigid Coupled Grasping
abstract
Grasping is fundamental to robotic manipulation, and recent advances in large-scale grasping datasets have provided essential training data and evaluation benchmarks, accelerating the development of learning-based methods for robust object grasping. However, most existing datasets exclude deformable bodies due to the lack of scalable, robust simulation pipelines, limiting the development of generalizable models for compliant grippers and soft manipulands. To address these challenges, we present GRIP, a General Robotic Incremental Potential contact simulation dataset for universal grasping. GRIP leverages an optimized Incremental Potential Contact (IPC)-based simulator for multi-environment data generation, achieving up to 48× speedup while ensuring efficient, intersection- and inversion-free simulations for compliant grippers and deformable objects. Our fully automated pipeline generates and evaluates diverse grasp interactions across 1,200 objects and 100,000 grasp poses, incorporating both soft and rigid grippers. The GRIP dataset enables applications such as neural grasp generation and stress field prediction. We release GRIP to advance research in robotic manipulation, soft-gripper control, and physics-driven simulation at: https://bell0o.github.io/GRIP/.
Siyu Ma, Wenxin Du, Chang Yu 0005, Zeshun Zong, Tianyi Xie, Yunuo Chen 0001, Yin Yang 0002, Xuchen Han, Chenfanfu Jiang
IROS6
2025 Negative label-Aware and correlation-Enhanced multi-Label feature selection
Huimin Fu 0002, Xiaoou Huang, Tianyi Xie, Lingfei Ren, Wanfu Gao, Yonghao Li, Xin Yang 0012
Knowl. Based Syst.4
2025 Dress-1-to-3: Single Image to Simulation-Ready 3D Outfit with Diffusion Prior and Differentiable Physics
abstract
Recent advances in large models have significantly advanced image-to-3D reconstruction. However, the generated models are often fused into a single piece, limiting their applicability in downstream tasks. This paper focuses on 3D garment generation, a key area for applications like virtual try-on with dynamic garment animations, which require garments to be separable and simulation-ready. We introduce Dress-1-to-3, a novel pipeline that reconstructs physics-plausible, simulation-ready separated garments with sewing patterns and humans from an in-the-wild image. Starting with the image, our approach combines a pre-trained image-to-sewing pattern generation model for creating coarse sewing patterns with a pre-trained multi-view diffusion model to produce multi-view images. The sewing pattern is further refined using a differentiable garment simulator based on the generated multi-view images. Versatile experiments demonstrate that our optimization approach substantially enhances the geometric alignment of the reconstructed 3D garments and humans with the input image. Furthermore, by integrating a texture generation module and a human motion generation module, we produce customized physics-plausible and realistic dynamic garment demonstrations. Our project page is https://dress-1-to-3.github.io/.
Xuan Li 0015, Chang Yu 0005, Wenxin Du, Tianyi Xie, Yunuo Chen 0001, Yin Yang 0002, Chenfanfu Jiang
ACM Trans. Graph.5
2024 PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
abstract
We introduce PhysGaussian, a new method that seamlessly integrates physically grounded Newtonian dynamics within 3D Gaussians to achieve high-quality novel motion synthe-sis. Employing a custom Material Point Method (MPM), our approach enriches 3D Gaussian kernels with physically meaningful kinematic deformation and mechanical stress attributes, all evolved in line with continuum mechanics principles. A defining characteristic of our method is the seamless integration between physical simulation and visual rendering: both components utilize the same 3D Gaussian kernels as their discrete representations. This negates the necessity for triangle/tetrahedron meshing, marching cubes, “cage meshes,” or any other geometry embedding, highlighting the principle of “what you see is what you simulate (WS2).” Our method demonstrates exceptional versatility across a wide variety of materials-including elastic entities, plastic metals, non-Newtonian fluids, and granular materials-showcasing its strong capabilities in creating diverse visual content with novel viewpoints and movements. Our project page is at: https://xpandora.github.io/PhysGaussian/
Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li 0015, Yutao Feng, Yin Yang 0002, Chenfanfu Jiang
CVPR1
2024 Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication
abstract
Existing diffusion-based text-to-3D generation methods primarily focus on producing visually realistic shapes and appearances, often neglecting the physical constraints necessary for downstream tasks. Generated models frequently fail to maintain balance when placed in physics-based simulations or 3D printed. This balance is crucial for satisfying user design intentions in interactive gaming, embodied AI, and robotics, where stable models are needed for reliable interaction. Additionally, stable models ensure that 3D-printed objects, such as figurines for home decoration, can stand on their own without requiring additional supports. To fill this gap, we introduce Atlas3D, an automatic and easy-to-implement method that enhances existing Score Distillation Sampling (SDS)-based text-to-3D tools. Atlas3D ensures the generation of self-supporting 3D models that adhere to physical laws of stability under gravity, contact, and friction. Our approach combines a novel differentiable simulation-based loss function with physically inspired regularization, serving as either a refinement or a post-processing module for existing frameworks. We verify Atlas3D's efficacy through extensive generation tasks and validate the resulting 3D models in both simulated and real-world environments.
Yunuo Chen 0001, Tianyi Xie, Zeshun Zong, Xuan Li 0015, Feng Gao 0013, Yin Yang 0002, Ying Nian Wu, Chenfanfu Jiang
NeurIPS2
2023 A Contact Proxy Splitting Method for Lagrangian Solid-Fluid Coupling
abstract
We present a robust and efficient method for simulating Lagrangian solid-fluid coupling based on a new operator splitting strategy. We use variational formulations to approximate fluid properties and solid-fluid interactions, and introduce a unified two-way coupling formulation for SPH fluids and FEM solids using interior point barrier-based frictional contact. We split the resulting optimization problem into a fluid phase and a solid-coupling phase using a novel time-splitting approach with augmented contact proxies , and propose efficient custom linear solvers. Our technique accounts for fluids interaction with nonlinear hyperelastic objects of different geometries and codimensions, while maintaining an algorithmically guaranteed non-penetrating criterion. Comprehensive benchmarks and experiments demonstrate the efficacy of our method.
Tianyi Xie, Minchen Li, Yin Yang 0002, Chenfanfu Jiang
ACM Trans. Graph.1
2022 Modeling human-human interaction with attention-based high-order GCN for trajectory prediction
Yanyan Fang, Zhiyu Jin, Zhenhua Cui, Qiaowen Yang, Tianyi Xie, Bo Hu 0002
Vis. Comput.5
2021 Towards Realistic Visual Dubbing with Heterogeneous Sources
abstract
The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data sources of videos and audios, thus causing the failure to leverage heterogeneous data sufficiently. In practice, it may be intractable to collect the perfect homologous data in some cases, for example, audio-corrupted or picture-blurry videos. To explore this kind of data and support high-fidelity few-shot visual dubbing, in this paper, we novelly propose a simple yet efficient two-stage framework with a higher flexibility of mining heterogeneous data. Specifically, our two-stage paradigm employs facial landmarks as intermediate prior of latent representations and disentangles the lip movements prediction from the core task of realistic talking head generation. By this means, our method makes it possible to independently utilize the training corpus for two-stage sub-networks using more available heterogeneous data easily acquired. Besides, thanks to the disentanglement, our framework allows a further fine-tuning for a given talking head, thereby leading to better speaker-identity preserving in the final synthesized results. Moreover, the proposed method can also transfer appearance features from others to the target speaker. Extensive experimental results demonstrate the superiority of our proposed method in generating highly realistic videos synchronized with the speech over the state-of-the-art.
Tianyi Xie, Liucheng Liao, Benlai Tang, Xiang Yin 0006, Jianfei Yang 0001, Jiali Yao, Yang Zhang 0088, Zejun Ma 0001
ACM Multimedia1
2021 CHECKER: Detecting Clickbait Thumbnails with Weak Supervision and Co-teaching
Tianyi Xie, Thai Le, Dongwon Lee 0001
ECML/PKDD (5)1
2020 Predicting Neural Deterioration in Patients with Alzheimer's Disease Using a Convolutional Neural Network
abstract
Alzheimer's disease causes neural damage, including brain atrophy in the patient. Consequently, ventricles that contain cerebral fluid a re e xpanded to filling th oseregions, which increases the proportional volume of ventricles in the brain. Therefore, abnormal growth of ventricle volume is an important indicator for estimating neural damage and, in turn, for the progression of Alzheimer's diseases. The rate of this volumegrowth, i.e., neural damage, can be predicted by predictive and machine learning models using the patient's current status. These predictions help assess the effectiveness of a particular treatment for a patient, in addition to providing some expectation of the disease timeline. In this work, we propose a convolutional neural network (CNN) model using region-level features for predicting ventricle volume biomarkers. The region-level representation with domaindriven features benefits from the CNN spatial pattern recognition capability. It also prevents learning irrelevant features and overfitting tot he t raining d ata a s a r esult 0 fd ata scarcity. Our model is applied to the ADNI dataset in the TADPOLE competition and outperforms the best leaderboard results.
Tavakoli H. Maryam, Tianyi Xie, Mirsad Hadzikadic, Yaorong Ge
BIBM2
2019 Dose Prediction for Prostate Radiation Treatment: Feasibility of a Distance-Based Deep Learning Model
abstract
This study aims to demonstrate the feasibility of using a novel distance-based representation of 3D CT-scan images to train a deep learning model for dose predictions in radiation treatment planning. The distance representation is inspired by previous knowledge of the domain to increase the generalizability of the deep learning models for radiation treatment planning. Conventional knowledge-based planning methods extract engineered features from 3D CT-scan images, as well as other patients' features, to predict the best achievable dose in a cancerous area and other organs at risk. Recent studies have shown higher accuracy in voxel-level dose prediction using deep learning models compared to the conventional machine learning approaches. Since the data resources for training these models are limited, most of the studies use 2D contour information to represent the patient anatomy. This representation loses volumetric information, and it is sensitive to small changes in patient orientation and translation. The distance-based representation introduced in this paper is inspired by the domain knowledge and is able to maintain the volumetric distance information despite the 2D slicing of 3D CT-image. According to prior studies in the radiation treatment planning domain, there is a strong association between the organs-at-risk distance from the cancerous volume and the patient's vulnerability to receive excessive dose. Therefore, the contour value in prior representation was replaced by voxel distance from cancerous volume. This modification in representation makes it transition and orientation invariant and adds potential robustness to patient positioning differences during the imaging/planning process. We evaluated the distance-based deep learning models through experiments for prediction of prostate cancer patients' vulnerability and voxel-level dose distribution using convolutional neural network and U-net models, respectively. The results were compared with contour-based U-net model as well as conventional machine learning with engineered representations. We found that the performance was comparable or higher than the prior state-of-the-art results for prostate-cancer dose distribution prediction.
Tavakoli H. Maryam, Boshu Ru, Tianyi Xie, Mirsad Hadzikadic, Qingrong Jackie Wu, Yaorong Ge
BIBM3
2018 Representing Knowledge for Radiation Therapy Planning with Markov Logic Networks
Yi Zhen, Tianyi Xie, Saugat Karki, Lulin Yuan, Qingrong Jackie Wu, Yaorong Ge
BIBM2
2018 Optimal Time Allocation in Relay Assisted Backscatter Communication Systems
abstract
In this paper, we consider a relay assisted backscatter communication (RaBackCom) system, where a user backscatters incident signals from a carrier emitter (CE) to a relay and a receiver simultaneously, and then the relay forwards the user's information to the receiver for throughput improvement. We consider two cases that the relay is with/without an embedded energy source. Specifically, if the relay does not have an energy source, it first harvests energy from the signals from the CE and then uses its harvested energy for information forwarding. For both cases, we formulate time allocation problems on the user's information backscattering, the user's information forwarding, or the relay's energy harvesting to maximize the system throughput, and then derive closed-form solutions. Simulation results demonstrate the advantages of the proposed relay cooperation scheme with the optimal time allocation in terms of system throughput.
Bin Lyu, Zhen Yang 0001, Tianyi Xie, Guan Gui 0001, Fumiyuki Adachi
VTC Spring3
2015 From Collision To Exploitation: Unleashing Use-After-Free Vulnerabilities in Linux Kernel
abstract
Since vulnerabilities in Linux kernel are on the increase, attackers have turned their interests into related exploitation techniques. However, compared with numerous researches on exploiting use-after-free vulnerabilities in the user applications, few efforts studied how to exploit use-after-free vulnerabilities in Linux kernel due to the difficulties that mainly come from the uncertainty of the kernel memory layout. Without specific information leakage, attackers could only conduct a blind memory overwriting strategy trying to corrupt the critical part of the kernel, for which the success rate is negligible.
Juanru Li, Junliang Shu, Tianyi Xie, Yuanyuan Zhang 0002, Dawu Gu
CCS5