Xi Zhao 0002

dblp:95/6322-2 · DBLP profile ↗
← Back
15ranked-venue papers
9as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 GSV-Pose: Pose estimation based on geometric similarity voting
Xi Zhao 0002, Yuekun Zhang, Jinji Wu
Appl. Intell.1
2024 Compress3D: A Compressed Latent Space for 3D Generation from a Single Image
Tianyu Yang 0003, Yu Li 0003, Lei Zhang 0001, Xi Zhao 0002
ECCV (18)5
2023 Reconstructing 3D Human Pose from RGB-D Data with Occlusions
abstract
Abstract We propose a new method to reconstruct the 3D human body from RGB‐D images with occlusions. The foremost challenge is the incompleteness of the RGB‐D data due to occlusions between the body and the environment, leading to implausible reconstructions that suffer from severe human‐scene penetration. To reconstruct a semantically and physically plausible human body, we propose to reduce the solution space based on scene information and prior knowledge. Our key idea is to constrain the solution space of the human body by considering the occluded body parts and visible body parts separately: modeling all plausible poses where the occluded body parts do not penetrate the scene, and constraining the visible body parts using depth data. Specifically, the first component is realized by a neural network that estimates the candidate region named the “free zone”, a region carved out of the open space within which it is safe to search for poses of the invisible body parts without concern for penetration. The second component constrains the visible body parts using the “truncated shadow volume” of the scanned body point cloud. Furthermore, we propose to use a volume matching strategy, which yields better performance than surface matching, to match the human body with the confined region. We conducted experiments on the PROX dataset, and the results demonstrate that our method produces more accurate and plausible results compared with other methods.
Bowen Dang, Xi Zhao 0002, He Wang 0002
Comput. Graph. Forum2
2023 Learning shape abstraction by cropping positive cuboid primitives with negative ones
Xi Zhao 0002
Vis. Comput.1
2022 Shape Completion with Points in the Shadow
abstract
Single-view point cloud completion aims to recover the full geometry of an object based on only limited observation, which is extremely hard due to the data sparsity and occlusion. The core challenge is to generate plausible geometries to fill the unobserved part of the object based on a partial scan, which is under-constrained and suffers from a huge solution space. Inspired by the classic shadow volume technique in computer graphics, we propose a new method to reduce the solution space effectively. Our method considers the camera a light source that casts rays toward the object. Such light rays build a reasonably constrained but sufficiently expressive basis for completion. The completion process is then formulated as a point displacement optimization problem. Points are initialized at the partial scan and then moved to their goal locations with two types of movements for each point: directional movements along the light rays and constrained local movement for shape refinement. We design neural networks to predict the ideal point movements to get the completion results. We demonstrate that our method is accurate, robust, and generalizable through exhaustive evaluation and comparison. Moreover, it outperforms state-of-the-art methods qualitatively and quantitatively on MVP datasets.
Xi Zhao 0002, He Wang 0002, Ruizhen Hu
SIGGRAPH Asia2
2022 Relationship-Based Point Cloud Completion
abstract
We propose a partial point cloud completion approach for scenes that are composed of multiple objects. We focus on pairwise scenes where two objects are in close proximity and are contextually related to each other, such as a chair tucked in a desk, a fruit in a basket, a hat on a hook and a flower in a vase. Different from existing point cloud completion methods, which mainly focus on single objects, we design a network that encodes not only the geometry of the individual shapes, but also the spatial relations between different objects. More specifically, we complete missing parts of the objects in a conditional manner, where the partial or completed point cloud of the other object is used as an additional input to help predict missing parts. Based on the idea of conditional completion, we further propose a two-path network, which is guided by a consistency loss between different sequences of completion. Our method can handle difficult cases where the objects heavily occlude each other. Also, it only requires a small set of training data to reconstruct the interaction area compared to existing completion approaches. We evaluate our method qualitatively and quantitatively via ablation studies and in comparison to the state-of-the-art point cloud completion methods.
Xi Zhao 0002, Jinji Wu, Ruizhen Hu, Taku Komura
IEEE Trans. Vis. Comput. Graph.1
2021 Learning Cuboid Abstraction of 3D Shapes via Iterative Error Feedback
Xi Zhao 0002, Yi-Jun Yang, Ruizhen Hu
Comput. Aided Des.1
2020 Informative scene decomposition for crowd analysis, comparison and simulation guidance
abstract
Crowd simulation is a central topic in several fields including graphics. To achieve high-fidelity simulations, data has been increasingly relied upon for analysis and simulation guidance. However, the information in real-world data is often noisy, mixed and unstructured, making it difficult for effective analysis, therefore has not been fully utilized. With the fast-growing volume of crowd data, such a bottleneck needs to be addressed. In this paper, we propose a new framework which comprehensively tackles this problem. It centers at an unsupervised method for analysis. The method takes as input raw and noisy data with highly mixed multi-dimensional (space, time and dynamics) information, and automatically structure it by learning the correlations among these dimensions. The dimensions together with their correlations fully describe the scene semantics which consists of recurring activity patterns in a scene, manifested as space flows with temporal and dynamics profiles. The effectiveness and robustness of the analysis have been tested on datasets with great variations in volume, duration, environment and crowd dynamics. Based on the analysis, new methods for data visualization, simulation evaluation and simulation guidance are also proposed. Together, our framework establishes a highly automated pipeline from raw data to crowd analysis, comparison and simulation guidance. Extensive experiments and evaluations have been conducted to show the flexibility, versatility and intuitiveness of our framework.
Feixiang He, Yuanhang Xiang, Xi Zhao 0002, He Wang 0002
ACM Trans. Graph.3
2020 Localization and Completion for 3D Object Interactions
abstract
Finding where and what objects to put into an existing scene is a common task for scene synthesis and robot/character motion planning. Existing frameworks require development of hand-crafted features suitable for the task, or full volumetric analysis that could be memory intensive and imprecise. In this paper, we propose a data-driven framework to discover a suitable location and then place the appropriate objects in a scene. Our approach is inspired by computer vision techniques for localizing objects in images: using an all directional depth image (ADD-image) that encodes the 360-degree field of view from samples in the scene, our system regresses the images to the positions where the new object can be located. Given several candidate areas around the host object in the scene, our system predicts the partner object whose geometry fits well to the host object. Our approach is highly parallel and memory efficient, and is especially suitable for handling interactions between large and small objects. We show examples where the system can hang bags on hooks, fit chairs in front of desks, put objects into shelves, insert flowers into vases, and put hangers onto laundry rack.
Xi Zhao 0002, Ruizhen Hu, Haisong Liu, Taku Komura, Xinyu Yang 0001
IEEE Trans. Vis. Comput. Graph.1
2020 Building hierarchical structures for 3D scenes with repeated elements
Xi Zhao 0002, Zhenqiang Su, Taku Komura, Xinyu Yang 0001
Vis. Comput.1
2019 Building hierarchical structures for 3D scenes based on normalized cut
abstract
Abstract The growing number of 3D scene data available online brings in new challenges for scene retrieval, understanding, and synthesis. Traditional shape processing methods have difficulty to manage 3D scenes because such methods ignore the contextual information, that is, the spatial relationship between the objects or groups of objects, which plays a significant role in describing scenes. Therefore, a context‐aware representation is needed to deal with such a problem. In this paper, we propose a method to build scene hierarchies based on contextual information. Given a 3D scene, we first use the interaction bisector surface to measure the affinity between different objects/elements of the scene and then apply the normalized cut method to build a hierarchical structure for the whole scene. The resulting hierarchical structure contains not only the relationship between the individual objects but also the relationship between object groups, which provides much richer information of the scene compared with a flat structure that only describes the contacts or affinity between the individual objects. We test our method using several public databases and show that the resulting structure is more consistent with the ground truth. We also show that our method can be used for point cloud segmentation and outperforms previous methods.
Xi Zhao 0002, Zhenqiang Su, Xinyu Yang 0001
Comput. Animat. Virtual Worlds1
2019 Regional classification of Chinese folk songs based on CRF model
Jing Luo 0007, Jianhang Ding, Xi Zhao 0002, Xinyu Yang 0001
Multim. Tools Appl.4
2019 Bidirectional Convolutional Recurrent Sparse Network (BCRSN): An Efficient Model for Music Emotion Recognition
abstract
Music emotion recognition, which enables effective and efficient music organization and retrieval, is a challenging subject in the field of music information retrieval. In this paper, we propose a new bidirectional convolutional recurrent sparse network (BCRSN) for music emotion recognition based on convolutional neural networks and recurrent neural networks. Our model adaptively learns the sequential-information-included affect-salient features (SII-ASF) from the 2-D time–frequency representation (i.e., spectrogram) of music audio signals. By combining feature extraction, ASF selection, and emotion prediction, the BCRSN can achieve continuous emotion prediction of audio files. To reduce the high computational complexity caused by the numerical-type ground truth, we propose a weighted hybrid binary representation (WHBR) method that converts the regression prediction process into a weighted combination of multiple binary classification problems. We test our method on two benchmark databases, that is, the Database for Emotional Analysis in Music and MoodSwings Turk. The results show that the WHBR method can greatly reduce the training time and improve the prediction accuracy. The extracted SII-ASF is robust to genre, timbre, and noise variation and is sensitive to emotion. It achieves significant improvement compared to the best performing feature sets in MediaEval 2015. Meanwhile, extensive experiments demonstrate that the proposed method outperforms the state-of-the-art methods.
Yizhuo Dong, Xinyu Yang 0001, Xi Zhao 0002
IEEE Trans. Multim.3
2016 Relationship templates for creating scene variations
abstract
We propose a novel example-based approach to synthesize scenes with complex relations, e.g., when one object is 'hooked', 'surrounded', 'contained' or 'tucked into' another object. Existing relationship descriptors used in automatic scene synthesis methods are based on contacts or relative vectors connecting the object centers. Such descriptors do not fully capture the geometry of spatial interactions, and therefore cannot describe complex relationships. Our idea is to enrich the description of spatial relations between object surfaces by encoding the geometry of the open space around objects, and use this as a template for fitting novel objects. To this end, we introduce relationship templates as descriptors of complex relationships; they are computed from an example scene and combine the interaction bisector surface (IBS) with a novel feature called the space coverage feature (SCF), which encodes the open space in the frequency domain. New variations of a scene can be synthesized efficiently by fitting novel objects to the template. Our method greatly enhances existing automatic scene synthesis approaches by allowing them to handle complex relationships, as validated by our user studies. The proposed method generalizes well, as it can form complex relationships with objects that have a topology and geometry very different from the example scene.
Xi Zhao 0002, Ruizhen Hu, Paul Guerrero 0001, Niloy J. Mitra, Taku Komura
ACM Trans. Graph.1
2014 Indexing 3D Scenes Using the Interaction Bisector Surface
abstract
The spatial relationship between different objects plays an important role in defining the context of scenes. Most previous 3D classification and retrieval methods take into account either the individual geometry of the objects or simple relationships between them such as the contacts or adjacencies. In this article we propose a new method for the classification and retrieval of 3D objects based on the Interaction Bisector Surface (IBS), a subset of the Voronoi diagram defined between objects. The IBS is a sophisticated representation that describes topological relationships such as whether an object is wrapped in, linked to, or tangled with others, as well as geometric relationships such as the distance between objects. We propose a hierarchical framework to index scenes by examining both the topological structure and the geometric attributes of the IBS. The topology-based indexing can compare spatial relations without being severely affected by local geometric details of the object. Geometric attributes can also be applied in comparing the precise way in which the objects are interacting with one another. Experimental results show that our method is effective at relationship classification and content-based relationship retrieval.
Xi Zhao 0002, He Wang 0002, Taku Komura
ACM Trans. Graph.1