Tiange Xiang

dblp:245/7663 · DBLP profile ↗
← Back
18ranked-venue papers
12as first author
16since 2021 · last 2026
0000-0003-3311-1287ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SocialGen: Modeling Multi-Human Social Interaction with Language Models
abstract
Human interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental to a wide range of real-world applications. In this paper, we introduce SocialGen, the first unified motionlanguage model capable of modeling interaction behaviors among varying numbers of individuals, to address this crucial yet challenging problem. Unlike prior methods that are limited to two-person interactions, we propose a novel social motion representation that supports tokenizing the motions of an arbitrary number of individuals and aligning them with the language space. This alignment enables the model to leverage rich, pretrained linguistic knowledge to better understand and reason about human social behaviors. To tackle the challenges of data scarcity, we curate a comprehensive multi-human interaction dataset, SocialX, enriched with textual annotations. Leveraging this dataset, we establish the first comprehensive benchmark for multihuman interaction tasks. Our method achieves state-of-theart performance across motion-language tasks, setting a new standard for multi-human interaction modeling. Our dataset and source code will be made publicly available.
Juze Zhang, Changan Chen, Tiange Xiang, Yusu Fang, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb
3DV4
2025 NeuHMR: Neural Rendering-Guided Human Motion Reconstruction
abstract
Reconstructing 3D human movements from video sequences is an important task in the fields of computer vision, graphics, and biomechanics. Although much progress has been made to infer 3D human mesh based on visual contexts provided in video sequences, generalization to in-the-wild videos still remains challenging for existing human mesh recovery (HMR) methods. To overcome inaccurate prediction, they can perform a second step optimization that refines the inaccurate estimations continuously at test time. Most optimization methods seek fitting of the body joints in the image space with respect to pseudo ground truth predicted by an off-the-shelf key point detector. However, state-of-theart detectors still introduce errors, especially for challenging poses. In this work, we rethink the dependency on the 2D key point fitting paradigm and present NeuHMR, an optimization-based mesh recovery framework based on recent advances in neural rendering. Our method builds on Human Neural Radiance Fields that allow the refinement of human meshes through animatable$2 D$renderings. We evaluated our method on two common benchmarks and validated its effectiveness.
Tiange Xiang, Kuan-Chieh Wang, Jaewoo Heo, Ehsan Adeli-Mosabbeb, Serena Yeung-Levy, Scott L. Delp, Li Fei-Fei 0001
3DV1
2025 Semantic Communications for Partially Observable Multi-Agent Reinforcement Learning-Based Unmanned Aerial Vehicles Monitoring System
Tiange Xiang, Seungwoo Seo, Sungwon Yi, Minseok Choi
ICC1
2025 Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation
abstract
Recent advances in text-to-image diffusion models have been driven by the increasing availability of paired 2D data. However, the development of 3D diffusion models has been hindered by the scarcity of high-quality 3D data, resulting in less competitive performance compared to their 2D counterparts. To address this challenge, we propose repurposing pre-trained 2D diffusion models for 3D object generation. We introduce Gaussian Atlas, a novel representation that utilizes dense 2D grids, enabling the fine-tuning of 2D diffusion models to generate 3D Gaussians. Our approach demonstrates successful transfer learning from a pre-trained 2D diffusion model to a 2D manifold flattened from 3D structures. To support model training, we compile GaussianVerse, a large-scale dataset comprising 205K high-quality 3D Gaussian fittings of various 3D objects. Our experimental results show that text-to-image diffusion models can be effectively adapted for 3D content generation, bridging the gap between 2D and 3D modeling.
Tiange Xiang, Chengjiang Long, Christian Häne, Peihong Guo, Scott L. Delp, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001
ICCV1
2025 MoDeFA: Multiobserver and Denoising-Enhanced Fingerprint Augmentation for Semi-Supervised Wi-Fi RSS-Based Indoor Positioning
abstract
Wi-Fi received signal strength (RSS) fingerprinting is a cost-effective indoor positioning method that has attracted broad attention because it only requires the existing Wi-Fi infrastructure and ubiquitous smart devices. However, conducting a site survey to establish a fingerprint database is time-consuming and labor-intensive, hindering the fast deployment of fingerprint-based indoor positioning systems. This paper presents a novel fingerprint reconstruction framework to reduce the site survey cost. The proposed MoDeFA is a semi-supervised deep learning framework that uses only a fraction of RSS data tagged with known locations to reconstruct the radio map. The resulting augmented fingerprints ensure the implementation of online fingerprint matching. First, we design a deep regression model to predict the RSS distribution from the perspectives of the multiple observational spots at surveyed locations. Next, we apply denoising to the initial radio map, a noisy representation of the ground truth, using the synthetic noise independent of the map errors without the need for ground-truth labels. Finally, we develop a noise-reduction model to compress the feature representation using unlabeled samples. Building on this, a model that combines multi-scale self-attention and gated scale selection can enhance the prediction of a target’s position. We tested MoDeFA on four public and a private dataset collected from the School of Software Engineering Building at Huazhong University of Science and Technology. The results indicate that the semi-supervised MoDeFA outperforms other state-of-the-art methods in various settings, even with only a fraction (30%) of the labeled RSS samples.
Tiange Xiang, Yuanjiang Sun
IEEE Internet Things J.1
2024 OccFusion: Rendering Occluded Humans with Generative Diffusion Priors
abstract
Existing human rendering methods require every part of the human to be fully visible throughout the input video. However, this assumption does not hold in real-life settings where obstructions are common, resulting in only partial visibility of the human. Considering this, we present OccFusion, an approach that utilizes efficient 3D Gaussian splatting supervised by pretrained 2D diffusion models for efficient and high-fidelity human rendering. We propose a pipeline consisting of three stages. In the Initialization stage, complete human masks are generated from partial visibility masks. In the Optimization stage, 3D human Gaussians are optimized with additional supervisions by Score-Distillation Sampling (SDS) to create a complete geometry of the human. Finally, in the Refinement stage, in-context inpainting is designed to further improve rendering quality on the less observed human body parts. We evaluate OccFusion on ZJU-MoCap and challenging OcMotion sequences and found that it achieves state-of-the-art performance in the rendering of occluded humans.
Adam Sun, Tiange Xiang, Scott L. Delp, Li Fei-Fei 0001, Ehsan Adeli-Mosabbeb
NeurIPS2
2024 Physical-space Multi-body Mesh Detection Achieved by Local Alignment and Global Dense Learning
abstract
From monocular RGB images captured in the wild, detecting multi-body 3D meshes in physical sizes and locations is notoriously difficult due to the diverse visual ambiguity and lack of explicit depth measurement. Modern DNN approaches made numerous advances based on either two-stage Region-of-Interests(RoI)-Align or single-stage fixed Field-of-View (FoV) detector frameworks for two main subtasks: local pelvis-centered mesh regression and global body-to-camera translation regression. However, sub-meter-level physical-space monocular mesh detection is still out of reach by existing solutions. In this paper, we recognize two common drawbacks: (1) The local meshes are usually estimated without explicitly aligning body features under image-space scaling, occlusion, and truncation; (2) The global translations are estimated based on a weak-perspective assumption, which tricks the network into prioritizing image-space (front-view) mesh alignment and leads to inaccurate mesh depth. We introduce Physical-space Multi-body Mesh Detection (PMMD), in which (1) Locally, we preserve the body aspect ratio, align the body-to-RoI layout, and densely refine the person-wise RoI features for robustness; (2) Globally, we learn dense-depth-guided features to amend the body-wise local feature for physical depth estimation. With the cleaned local features and explicit local-global associations, PMMD achieves the best centimeter-level local mesh metrics and the first sub-meter-level global mesh metrics from monocular images in 3DPW and AGORA datasets.
Haoye Dong, Tiange Xiang, Sravan Chittupalli, Jun Liu 0075
WACV2
2024 Exploiting Structural Consistency of Chest Anatomy for Unsupervised Anomaly Detection in Radiography Images
abstract
Radiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. Exploiting this structured information could potentially ease the detection of anomalies from radiography images. To this end, we propose a Simple Space-Aware Memory Matrix for In-painting and Detecting anomalies from radiography images (abbreviated as SimSID). We formulate anomaly detection as an image reconstruction task, consisting of a space-aware memory matrix and an in-painting block in the feature space. During the training, SimSID can taxonomize the ingrained anatomical structures into recurrent visual patterns, and in the inference, it can identify anomalies (unseen/modified visual patterns) from the test image. Our SimSID surpasses the state of the arts in unsupervised anomaly detection by +8.0%, +5.0%, and +9.9% AUC scores on ZhangLab, COVIDx, and CheXpert benchmark datasets, respectively.
Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision Decoding
abstract
Decoding visual stimuli from brain recordings aims to deepen our understanding of the human visual system and build a solid foundation for bridging human and computer vision through the Brain-Computer Interface. However, reconstructing high-quality images with correct semantics from brain recordings is a challenging problem due to the complex underlying representations of brain signals and the scarcity of data annotations. In this work, we present MinD-Vis: Sparse Masked Brain Modeling with Double-Conditioned Latent Diffusion Model for Human Vision Decoding. Firstly, we learn an effective self-supervised representation of fMRI data using mask modeling in a large latent space inspired by the sparse coding of information in the primary visual cortex. Then by augmenting a latent diffusion model with double-conditioning, we show that MinD-Vis can reconstruct highly plausible images with semantically matching details from brain recordings using very few paired annotations. We benchmarked our model qualitatively and quantitatively; the experimental results indicate that our method outperformed state-of-the-art in both semantic mapping (100-way semantic classification) and generation quality (FID) by 66% and 41% respectively. An exhaustive ablation study was also conducted to analyze our framework.
Zijiao Chen, Jiaxin Qing, Tiange Xiang, Wan Lin Yue, Juan Helen Zhou
CVPR3
2023 SQUID: Deep Feature In-Painting for Unsupervised Anomaly Detection
abstract
Radiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. To exploit this structured information, we propose the use of Space-aware Memory Queues for In-painting and Detecting anomalies from radiography images (abbreviated as SQUID). We show that SQUID can taxonomize the ingrained anatomical structures into recurrent patterns; and in the inference, it can identify anomalies (unseen/modified patterns) in the image. SQUID surpasses 13 state-of-the-art methods in unsupervised anomaly detection by at least 5 points on two chest X-ray benchmark datasets measured by the Area Under the Curve (AUC). Additionally, we have created a new dataset (DigitAnatomy), which synthesizes the spatial correlation and consistent shape in chest anatomy. We hope DigitAnatomy can prompt the development, evaluation, and interpretability of anomaly detection methods.
Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou
CVPR1
2023 Rendering Humans from Object-Occluded Monocular Videos
abstract
3D understanding and rendering of moving humans from monocular videos is a challenging task. Despite recent progress, the task remains difficult in real-world scenarios, where obstacles may block the camera view and cause partial occlusions in the captured videos. Existing methods cannot handle such defects due to two reasons. First, the standard rendering strategy relies on point-point mapping, which could lead to dramatic disparities between the visible and occluded areas of the body. Second, the naive direct regression approach does not consider any feasibility criteria (i.e., prior information) for rendering under occlusions. To tackle the above drawbacks, we present OccNeRF, a neural rendering method that achieves better rendering of humans in severely occluded scenes. As direct solutions to the two drawbacks, we propose surface-based rendering by integrating geometry and visibility priors. We validate our method on both simulated and real-world occlusions and demonstrate our method’s superiority. Project page: https://cs.stanford.edu/~xtiange/projects/occnerf/
Tiange Xiang, Adam Sun, Jiajun Wu 0001, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001
ICCV1
2023 DDM2: Self-Supervised Diffusion MRI Denoising with Generative Diffusion Models
Tiange Xiang, Mahmut Yurt, Ali B. Syed, Kawin Setsompop, Akshay Chaudhari
ICLR1
2022 Towards bi-directional skip connections in encoder-decoder architectures and beyond
Tiange Xiang, Chaoyi Zhang, Xinyi Wang 0015, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai
Medical Image Anal.1
2022 DSNet: A Dual-Stream Framework for Weakly-Supervised Gigapixel Pathology Image Analysis
abstract
We present a novel weakly-supervised framework for classifying whole slide images (WSIs). WSIs, due to their gigapixel resolution, are commonly processed by patch-wise classification with patch-level labels. However, patch-level labels require precise annotations, which is expensive and usually unavailable on clinical data. With image-level labels only, patch-wise classification would be sub-optimal due to inconsistency between the patch appearance and image-level label. To address this issue, we posit that WSI analysis can be effectively conducted by integrating information at both high magnification (local) and low magnification (regional) levels. We auto-encode the visual signals in each patch into a latent embedding vector representing local information, and down-sample the raw WSI to hardware-acceptable thumbnails representing regional information. The WSI label is then predicted with a Dual-Stream Network (DSNet), which takes the transformed local patch embeddings and multi-scale thumbnail images as inputs and can be trained by the image-level label only. Experiments conducted on three large-scale public datasets demonstrate that our method outperforms all recent state-of-the-art weakly-supervised WSI classification methods.
Tiange Xiang, Yang Song 0001, Chaoyi Zhang, Dongnan Liu, Fan Zhang 0013, Heng Huang 0001, Lauren O'Donnell, Tom Weidong Cai
IEEE Trans. Medical Imaging1
2021 Walk in the Cloud: Learning Curves for Point Clouds Shape Analysis
abstract
Discrete point cloud objects lack sufficient shape descriptors of 3D geometries. In this paper, we present a novel method for aggregating hypothetical curves in point clouds. Sequences of connected points (curves) are initially grouped by taking guided walks in the point clouds, and then subsequently aggregated back to augment their pointwise features. We provide an effective implementation of the proposed aggregation strategy including a novel curve grouping operator followed by a curve aggregation operator. Our method was benchmarked on several point cloud analysis tasks where we achieved the state-of-the-art classification accuracy of 94.2% on the ModelNet40 classification task, instance IoU of 86.8% on the ShapeNetPart segmentation task and cosine error of 0.11 on the ModelNet40 normal estimation task. Our project page with source code is available at: https://curvenet.github.io/.
Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Jianhui Yu, Tom Weidong Cai
ICCV1
2021 BiX-NAS: Searching Efficient Bi-directional Architecture for Medical Image Segmentation
Xinyi Wang 0015, Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai
MICCAI (1)2
2020 BiO-Net: Learning Recurrent Bi-directional Connections for Encoder-Decoder Architecture
Tiange Xiang, Chaoyi Zhang, Dongnan Liu, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai
MICCAI (1)1
2019 Integration of Multimodal Data for Breast Cancer Classification Using a Hybrid Deep Learning Method
Rui Yan 0009, Xiaosong Rao, Baorong Shi, Tiange Xiang, Chun-Hou Zheng 0001, Fa Zhang 0001
ICIC (1)5