Xin Zhao 0025

dblp:68/2766-25 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-6207-5401ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GVRRI: Identifying visual receptive regions in node-link diagrams for node-centered graph analysis
Xin Zhao 0025, Luanxi Huang, Ning Zhang 0007, Wenjian Zuo, Ying Zhao 0001
Comput. Graph.1
2026 EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion
abstract
Speech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies.
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, Hong Qin 0001, Yang Gao 0032
IEEE Trans. Vis. Comput. Graph.1
2026 EmoDiffuser: emotional diffuser for speech-driven 3D facial animation
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, JunJun Pan, Yang Gao 0032
Vis. Comput.1
2025 Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
abstract
In 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant coupling in self-supervised audio feature spaces, leading to the averaging effect in subsequent lip motion generation. To address this issue, this paper proposes a plug-and-play semantic decorrelation module—Wav2Sem. This module extracts semantic features corresponding to the entire audio sequence, leveraging the added semantic information to decorrelate audio encodings within the feature space, thereby achieving more expressive audio features. Extensive experiments across multiple Speech-driven models indicate that the Wav2Sem module effectively decouples audio features, significantly alleviating the averaging effect of phonetically similar syllables in lip shape generation, thereby enhancing the precision and naturalness of facial animations. Our source code is available at https://github.com/wslh852/Wav2Sem.git.
Ju Dai, Xin Zhao 0025, Feng Zhou 0007, JunJun Pan, Lei Li 0050
CVPR3
2025 TransportMap: Visual transport analysis for spatiotemporal data without trajectory information
Jiazhi Xia, Xin Zhao 0025, Kang Xie, Yangbo Hou, Xiaolong (luke) Zhang, Xiaoyan Kui, Ying Zhao 0001, Chenhui Li 0001, Hong Qin 0001
Comput. Graph.2
2025 An Open Dataset of Cyber Asset Graphs for Cybercrime Research
abstract
Cybercrime poses a severe threat to the entire Internet ecosystem. Various cyber assets, such as domain name, IP address, and security certificate, are staple infrastructures of cybercrime. A cyber asset graph (CAG) is a collection of closely related cyber assets held by a cybercrime gang to support online criminal activities. Analyzing CAGs provides rich data insights for cybercrime investigation and governance. This paper introduces an open dataset of CAGs comprised of 2.37 million nodes with eight types of cyber assets and 3.28 million edges with eleven types of relations. This paper introduces the dataset construction process, applied areas, and the experience of using the dataset in the ChinaVis Data Challenge 2022. This dataset contains numerous CAGs of cybercrime gangs in the real world, which is the first open dataset of CAGs for cybercrime research. This dataset can also support the development of other application-oriented areas, such as cyber asset management and cyber-physical-social system, and various graph-related research areas, such as graph theory, graph mining, and graph visualization.
Xin Zhao 0025, Shaolong Li, Ying Zhao 0001, Shuowen Fu, Yunpeng Chen, Zhuo Chen 0029
IEEE Trans. Big Data1
2025 Investigating Visual Perception of Degree Centrality in Graph Visualization
abstract
Degree centrality (DC) is a widely used metric that measures node importance in data space. A node-link diagram is a commonly used graph visualization to help viewers identify important nodes in visual space. Previous graph perception studies largely concentrated on revealing perception principles in visual space. However, they rarely investigated the intrinsic relations between computed and perceived important nodes by jointly using data and visual spaces, thereby hindering a deep integration of computational and interactive graph analytics. To address this gap, we adopted the visual perception of DC as a representative object to conduct a graph perception study by jointly using data and visual spaces. Two research questions were defined. (RQ1) Can viewers accurately estimate the relative DCs of the given nodes in a node-link diagram through visual perception? (RQ2) What visual factors influence viewers' visual estimation of relative DCs? A controlled user experiment was conducted to answer the questions. Results showed that: (1) The participants failed to estimate the relative DCs accurately, particularly when the DC differences between nodes were not great. (2) Seven visual factors influencing the tasks were summarized, such as the size of the visual receptive region of a node, the link and node densities in the visual receptive region of the node, and the wrapping angle of the node's neighbors. (3) The factors presented certain priorities in complicated situations. These findings provide rich implications for graph analytics, such as utilizing the findings to optimize graph visualizations to achieve the desired consistency between computed and perceived important nodes.
Xin Zhao 0025, Shuowen Fu, Yunpeng Chen, Ying Zhao 0001
IEEE Trans. Vis. Comput. Graph.1
2023 Interactive Visual Cluster Analysis by Contrastive Dimensionality Reduction
abstract
We propose a contrastive dimensionality reduction approach (CDR) for interactive visual cluster analysis. Although dimensionality reduction of high-dimensional data is widely used in visual cluster analysis in conjunction with scatterplots, there are several limitations on effective visual cluster analysis. First, it is non-trivial for an embedding to present clear visual cluster separation when keeping neighborhood structures. Second, as cluster analysis is a subjective task, user steering is required. However, it is also non-trivial to enable interactions in dimensionality reduction. To tackle these problems, we introduce contrastive learning into dimensionality reduction for high-quality embedding. We then redefine the gradient of the loss function to the negative pairs to enhance the visual cluster separation of embedding results. Based on the contrastive learning scheme, we employ link-based interactions to steer embeddings. After that, we implement a prototype visual interface that integrates the proposed algorithms and a set of visualizations. Quantitative experiments demonstrate that CDR outperforms existing techniques in terms of preserving correct neighborhood structures and improving visual cluster separation. The ablation experiment demonstrates the effectiveness of gradient redefinition. The user study verifies that CDR outperforms t-SNE and UMAP in the task of cluster identification. We also showcase two use cases on real-world datasets to present the effectiveness of link-based interactions.
Jiazhi Xia, Linquan Huang, Weixing Lin, Xin Zhao 0025, Jing Wu 0004, Yang Chen 0048, Ying Zhao 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.4
2022 Evaluating Effects of Background Stories on Graph Perception
abstract
A graph is an abstract model that represents relations among entities, for example, the interactions between characters in a novel. A background story endows entities and relations with real-world meanings and describes the semantics and context of the abstract model, for example, the actual story that the novel presents. Considering practical experience and prior research, human viewers who are familiar with the background story of a graph and those who do not know the background story may perceive the same graph differently. However, no previous research has adequately addressed this problem. This research article thus presents an evaluation that investigated the effects of background stories on graph perception. Three hypotheses that focused on the role of visual focus areas, graph structure identification, and mental model formation on graph perception were formulated and guided three controlled experiments that evaluated the hypotheses using real-world graphs with background stories. An analysis of the resulting experimental data, which compared the performance of participants who read and did not read the background stories, obtained a set of instructive findings. First, having knowledge about a graph's background story influences participants' focus areas during interactive graph explorations. Second, such knowledge significantly affects one's ability to identify community structures but not high degree and bridge structures. Third, this knowledge influences graph recognition under blurred visual conditions. These findings can bring new considerations to the design of storytelling visualizations and interactive graph explorations.
Ying Zhao 0001, Jingcheng Shi, Jiawei Liu 0001, Jian Zhao 0010, Wenzhi Zhang, Kangyi Chen, Xin Zhao 0025, Chunyao Zhu, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.8
2021 An Indoor Crowd Movement Trajectory Benchmark Dataset
abstract
In recent years, technologies of indoor crowd positioning and movement data analysis have received widespread attention in the fields of reliability management, indoor navigation, and crowd behavior monitoring. However, only a few indoor crowd movement trajectory datasets are available to the public, thus restricting the development of related research and application. This article contributes a new benchmark dataset of indoor crowd movement trajectories. This dataset records the movements of over 5000 participants at a three-day large academic conference in a two-story indoor venue. The conference comprises varied activities, such as academic seminars, business exhibitions, a hacking contest, interviews, tea breaks, and a banquet. The participants are divided into seven types according to participation permission to the activities. Some of them are involved in anomalous events, such as loss of items, unauthorized accesses, and equipment failures, forming a variety of spatial–temporal movement patterns. In this article, we first introduce the scenario design, entity and behavior modeling, and data generator of the dataset. Then, a detailed ground truth of the dataset is presented. Finally, we describe the process and experience of applying the dataset to the contest of ChinaVis Data Challenge 2019. Evaluation results of the 75 contest entries and the feedback from 359 contestants demonstrate that the dataset has satisfactory completeness, and usability, and can effectively identify the performance of methods, technologies, and systems for indoor trajectory analysis.
Ying Zhao 0001, Xin Zhao 0025, Siming Chen 0001
IEEE Trans. Reliab.2
2015 Real-time haptic manipulation and cutting of hybrid soft tissue models by extended position-based dynamics
abstract
Abstract This paper systematically describes an interactive dissection approach for hybrid soft tissue models governed by extended position‐based dynamics. Our framework makes use of a hybrid geometric model comprising both surface and volumetric meshes. The fine surface triangular mesh with high‐precision geometric structure and texture at the detailed level is employed to represent the exterior structure of soft tissue models. Meanwhile, the interior structure of soft tissues is constructed by coarser tetrahedral mesh, which is also employed as physical model participating in dynamic simulation. The less details of interior structure can effectively reduce the computational cost during simulation. For physical deformation, we design and implement an extended position‐based dynamics approach that supports topology modification and material heterogeneities of soft tissue. Besides stretching and volume conservation constraints, it enforces the energy preserving constraints, which take the different spring stiffness of material into account and improve the visual performance of soft tissue deformation. Furthermore, we develop mechanical modeling of dissection behavior and analyze the system stability. The experimental results have shown that our approach affords real‐time and robust cutting without sacrificing realistic visual performance. Our novel dissection technique has already been integrated into a virtual reality‐based laparoscopic surgery simulator. Copyright © 2015 John Wiley & Sons, Ltd.
JunJun Pan, Junxuan Bai, Xin Zhao 0025, Aimin Hao, Hong Qin 0001
Comput. Animat. Virtual Worlds3
2015 Metaballs-based physical modeling and deformation of organs for virtual surgery
JunJun Pan, Chengkai Zhao, Xin Zhao 0025, Aimin Hao, Hong Qin 0001
Vis. Comput.3
2014 Dissection of hybrid soft tissue models using position-based dynamics
abstract
This paper describes an interactive dissection approach for hybrid soft tissue models governed by position-based dynamics. Our framework makes use of a hybrid geometric model comprising both surface and volumetric meshes. The fine surface triangular mesh is used to represent the exterior structure of soft tissue models. Meanwhile, the interior structure of soft tissues is constructed by coarser tetrahedral meshes, which are also employed as physical models participating in dynamic simulation. The less details of interior structure can effectively reduce the computational cost of deformation and geometric subdivision during dissection. For physical deformation, we design and implement a position-based dynamics approach that supports topology modification and enforces the volume-preserving constraint. Experimental results have shown that, this hybrid dissection method affords real-time and robust cutting simulation without sacrificing realistic visual performance.
JunJun Pan, Junxuan Bai, Xin Zhao 0025, Aimin Hao, Hong Qin 0001
VRST3
2013 Four-Dimensional Geometry Lens: A Novel Volumetric Magnification Approach
abstract
Abstract We present a novel methodology that utilizes four‐dimensional (4D) space deformation to simulate a magnification lens on versatile volume datasets and textured solid models. Compared with other magnification methods (e.g. geometric optics, mesh editing), 4D differential geometry theory and its practices are much more flexible and powerful for preserving shape features (i.e. minimizing angle distortion), and easier to adapt to versatile solid models. The primary advantage of 4D space lies at the following fact: we can now easily magnify the volume of regions of interest (ROIs) from the additional dimension, while keeping the rest region unchanged. To achieve this primary goal, we first embed a 3D volumetric input into 4D space and magnify ROIs in the fourth dimension. Then we flatten the 4D shape back into 3D space to accommodate other typical applications in the real 3D world. In order to enforce distortion minimization, in both steps we devise the high‐dimensional geometry techniques based on rigorous 4D geometry theory for 3D/4D mapping back and forth to amend the distortion. Our system can preserve not only focus region, but also context region and global shape. We demonstrate the effectiveness, robustness and efficacy of our framework with a variety of models ranging from tetrahedral meshes to volume datasets.
Bo Li 0014, Xin Zhao 0025, Hong Qin 0001
Comput. Graph. Forum2