Jiayi Xu 0002

dblp:40/6214-2 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-9868-2913ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 BeautyMark: A diffusion model for aesthetic QR code generation with robust watermark authentication
Jiayi Xu 0002, Hangpeng Ren, Jianfeng Lu 0005, Li Li 0014, Mahmoud Emam
Appl. Intell.1
2026 MSGS: Multi-space Gaussian Splatting for mirror reflections
Zhankong Bao, Jiayi Xu 0002, Xuanxuan Huang, Masahiro Toyoura, Gefei Xie, Qianhong Xiang, Huaming Lin
Comput. Graph.2
2025 Feature Disentanglement in GANs for Photorealistic Multi-view Hair Transfer
abstract
Abstract Fast and highly realistic multi‐view hair transfer plays a crucial role in evaluating the effectiveness of virtual hair try‐on systems. However, GAN‐based generation and editing methods face persistent challenges in feature disentanglement. Achieving pixel‐level, attribute‐specific modifications—such as changing hairstyle or hair color without affecting other facial features—remains a long‐standing problem. To address this limitation, we propose a novel multi‐view hair transfer framework that leverages a hair‐only intermediate facial representation and a 3D‐guided masking mechanism. Our approach disentangles tri‐plane facial features into spatial geometric components and global style descriptors, enabling independent and precise control over hairstyle and hair color. By introducing a dedicated intermediate representation focused solely on hair and incorporating a two‐stage feature fusion strategy guided by the generated 3D mask, our framework achieves fine‐grained local editing across multiple viewpoints while preserving facial integrity and improving background consistency. Extensive experiments demonstrate that our method produces visually compelling and natural results in side‐to‐front view hair transfer tasks, offering a robust and flexible solution for high‐fidelity hair reconstruction and manipulation.
Jiayi Xu 0002, Chenming Zhang, Xiaogang Jin 0001, Yaohua Ji
Comput. Graph. Forum1
2025 High similarity controllable face anonymization based on dynamic identity perception
Jiayi Xu 0002, Yixuan Ju, Xiaoyang Mao, Shanqing Zhang
Vis. Comput.1
2024 Visual Coherence Face Anonymization Algorithm Based on Dynamic Identity Perception
abstract
In the era of the meta-universe and the proliferation of personalized social networks, interactive behaviors like sharing personal and family photos pose an escalating risk of privacy breaches and identity exposure. A potential remedy lies in substituting real images with anonymized face images in public contexts. While existing face anonymization methods often replace substantial portions of face images, the resultant faces lack sufficient similarity to the originals. To address this, we propose a anonymization model leveraging saliency analysis to detect identity relevant facial region, preserving visual coherence and avoiding recognition by face recognition systems. Our model comprises two integral networks: the Dynamic Identity Perception Network (DIPNet) and the improved PSPNet. DIPNet, in particular, encompasses two vital sub-modules: the dynamic region perception module detects identity relevant region; the anonymization region control module governs the size of region through thresholding, thereby dominating the preservation of identity independent features and the degree of anonymization. The improved PSPNet produces high-quality identity anonymized faces. Experimental results demonstrate that our method yields realistic anonymized faces, retaining original features and deceiving face recognition systems, safeguarding privacy in the modern digital landscape.
Shanqing Zhang, Yixuan Ju, Xiaoyang Mao, Jiayi Xu 0002
FG5
2024 Action recognition algorithm based on skeleton graph with multiple features and improved adjacency matrix
abstract
Abstract Although graph convolutional networks have achieved good performances in skeleton‐graph‐based action recognition, there are still some problems which include the incomplete utilization of skeleton graph features and the lacking of logical adjacency information between nodes in adjacency matrix. In this article, a human action recognition algorithm is proposed based on multiple features from the skeleton graph to solve these problems. More specifically, an improved adjacency matrix is constructed to make full use of the multiple skeleton graph features. These features include local differential features, multi‐scale edge features, features of the original skeleton graph, nodal features, and nodal motion features. Extensive results are conducted on four standard datasets (NTU RGB‐D 60, NTU RGB‐D 120, Kinetics, and Northwestern‐UCLA). The experimental results show that the proposed algorithm outperforms the SOTA action recognition algorithms.
Shanqing Zhang, Shuheng Jiao, Jiayi Xu 0002
IET Image Process.4
2024 Personalized hairstyle and hair color editing based on multi-feature fusion
abstract
Abstract In the metaverse era, virtual design of hairstyle becomes very popular for personalized aesthetics. As hair design tasks can be decomposed into hair attribute editing and generation, the development of generative adversarial networks (GANs) has significantly prompted its development. The majority of the existing algorithms focus on transferring the overall hair region from one face to another, which ignore fine control over the color and geometric features. Furthermore, these algorithms may result in unnatural generation results. In this paper, we propose a hair modification framework that learns hairstyle information from a reference face mask and color information from a guidance face image. Firstly, the features of the input face image and reference images are extracted through a group of encoders, and then divided into feature vectors of coarse, medium, and fine levels. Secondly, multi-level feature vectors are fused in the latent space using attention-based modulation modules. Finally, the fused feature vector is passed through a StyleGAN generator to generate face images with specified hairstyle and hair color. Experimental results show that the proposed method can finely simulate the hairstyle transition between long and short hair under the constraint of the reference mask, and can produce realistic fusion effects in the hair-covered regions, such as ears, neck, and forehead. Various hair dyeing effects that adapt to personalized characteristics are demonstrated, as facial features including skin color and hair texture are preserved when transferring the hair color.
Jiayi Xu 0002, Chenming Zhang, Weikang Zhu, Li Li 0014, Xiaoyang Mao
Vis. Comput.1
2021 Adaptive semantic attribute decoupling for precise face image editing
Yixuan Ju, Xiaoyang Mao, Jiayi Xu 0002
Vis. Comput.4
2021 Matching a composite sketch to a photographed face using fused HOG and deep feature models
Jiayi Xu 0002, Xinying Xue, Yitiao Wu, Xiaoyang Mao
Vis. Comput.1
2020 Person-independent facial expression recognition method based on improved Wasserstein generative adversarial networks in combination with identity aware
Caie Xu, Yunhui Zhang, Jiayi Xu 0002
Multim. Syst.5
2020 Image enhancement algorithm based on generative adversarial network in combination of improved game adversarial loss mechanism
Caie Xu, Yunhui Zhang, Jiayi Xu 0002
Multim. Tools Appl.5
2019 Composite Sketch Recognition Using Multi-scale Hog Features and Semantic Attributes
abstract
Composite sketch recognition belongs to heterogeneous face recognition research, which is of great important in the field of criminal investigation. Because composite face sketch and photo belong to different modalities, robust representation of face feature cross different modalities is the key to recognition. Considering that composite sketch lacks texture details in some area, using texture features only may result in low recognition accuracy, this paper proposes a composite sketch recognition algorithm based on multi-scale Hog features and semantic attributes. Firstly, the global Hog features of the face and the local Hog features of each face component are extracted to represent the contour and detail features. Then the global and detail features are fused according to their importance at score level. Finally, semantic attributes are employed to reorder the matching results. The proposed algorithm is validated on PRIP-VSGC database and UoM-SGFS database, and achieves rank 10 identification accuracy of 88.6% and 96.7% respectively, which demonstrates that the proposed method outperforms other state-of-the-art methods.
Xinying Xue, Jiayi Xu 0002, Xiaoyang Mao
CW2
2018 Suggesting the Appropriate Number of Observers for Predicting Video Saliency with Eye-Tracking Data
abstract
Accurately predicting video saliency is important for applications such as video quality assessment, summary, compression, and retargeting. As the automatic saliency models for videos suffer from problems of inaccuracy, determining video saliency from data on the human gaze is a promising approach. Due to differences in individual observers, however, eye-tracking data of a certain number of observers are usually required to compute a visual attention map close to the ground truth. Although it has become cheaper to acquire human eye-tracking data thanks to the lower price of equipment, it is still not easy to carry out studies with a large number of observers. To keep the balance between accuracy and expense, this paper proposes a new method for suggesting the appropriate number of observers needed in eye-tracking experiments for a given video. Through carefully analyzing eye-tracking data of various video clips, we found videos can be classified into four types based on the number of observers required to approach the ground truth. A new support vector machine (SVM) classifier was trained to automatically classify videos into one of the four typical types.
Chuancai Li, Jiayi Xu 0002, Jianjun Li 0001, Xiaoyang Mao
CGI2
2017 Synthesis of Facial Images Based on Relevance Feedback
abstract
We propose a dialogic system based on a relevance feedback strategy that allows for the semiautomatic synthesis of a facial image that only exists in a user's mind. The user is presented with several facial images and judges whether each one resembles the face that he or she is imagining. Based on the feedback from the user, a set of sample facial images are used to train an Optimum-Path Forest classifying the relevance of facial images. An interpolation method is then employed to synthesize new facial images that closely resemble the imagined face. A series of experiments are conducted to evaluate and verify the effectiveness and efficiency of the proposed technique.
Caie Xu, Shota Fushimi, Masahiro Toyoura, Jiayi Xu 0002, Xiaoyang Mao
CW4
2016 Example-based caricature generation with exaggeration control
Masahiro Toyoura, Jiayi Xu 0002, Fumio Ohnuma, Xiaoyang Mao
Vis. Comput.3
2015 Hidden message in a deformation-based texture
Jiayi Xu 0002, Xiaoyang Mao, Xiaogang Jin 0001, Aubrey Jaffer, Shufang Lu, Li Li 0014, Masahiro Toyoura
Vis. Comput.1
2014 A Study on Perceived Similarity between Photograph and Shape Exaggerated Caricature
abstract
This paper investigates the relationship between the extent of exaggeration in a caricature and its face identification ability. As face recognition is largely influenced by facial deformations, we focused on finding the borderline between likeness and unlikeness by applying gradual alterations to the face shape of the subject being studied. Suggestions on manipulating the degree of similarity when generating a caricature will be given. The experimental environment in this research can be used as a user-friendly caricature generation system based on Exaggerating the Difference From the Mean face, which allows a user to freely control each generation step and design his or her own unique caricature portrait.
Jiayi Xu 0002, Xiaoyang Mao, Masahiro Toyoura, Xiaogang Jin 0001
CW1
2014 Example-Based Automatic Caricature Generation
abstract
Caricature is a popular artistic media widely used for effective communications. The fascination of caricature lies in its expressive depiction of a person's prominent features, which is usually realized through the so called exaggeration technique. This paper proposes a new example based automatic caricature generation system supporting the exaggeration of visual appearance features. The system comprises the construction of a learning database and the generation of caricatures. The construction of the learning database links the pairs of facial images and corresponding caricatures. Given an input face, the system automatically compute the feature vectors of facial parts and hairstyle, and search the learning database for the exaggerated parts by using the most prominent features. Experimental results show that our system can achieve the control over the degree of exaggeration and the exaggerated results can better represent the features of the subjects.
Kouki Tajima, Jiayi Xu 0002, Masahiro Toyoura, Xiaoyang Mao
CW3
2013 Stego-Marbling-Texture
abstract
We present stego-marbling-texture, a new and unique texture design method which allows users to deliver personalized messages with beautiful marbling textures. Our approach is inspired by the success of the recent work on modeling traditional marbling operations as mathematical functions. The encrypter transforms an input image or a text message into an intricate marbling pattern using marbling operations defined as reversible functions, and the decrypter recovers the input image or message through reversing the process of marbling operations. When applying marbling operations, the parameters of operations are automatically recorded, encrypted, and then invisibly embedded into the marbling pattern to create a stego-marbling-texture. In this way, the decrypter can be implemented as a stand along software, enabling the receiver to extract the hidden message from the stego-marbling-texture without requiring any extra information from the sender. To ensure that the message is unnoticeably and beautifully covered by the marbling texture, we propose a new technique for automatically creating a background which is harmonious with the input message based on a set of visual perception cues.
Jiayi Xu 0002, Xiaoyang Mao, Xiaogang Jin 0001, Aubrey Jaffer, Shufang Lu, Li Li 0014, Masahiro Toyoura
CAD/Graphics1
2008 Shape-constrained flock animation
abstract
Abstract We propose a novel shape‐constrained flock animation system for interactively controlling flock navigation in virtual environments. This system is capable of making the spatial distribution of a flock meet static or deforming shape constraints while performing flock simulation. Such a capability can find many applications in the entertainment industry. Given a 3D constraining shape, our system first draws a set of uniform sample points through a 3D surface mosaicing process or a stratified point sampling strategy. Once correspondences between flock members and sample points have been established, points on the target shape are used as homing destinations to guide flock migration. Under a global path control scheme, an effective fuzzy control logic, which dynamically adjusts steering forces and control forces, has been developed to create visually pleasing shape‐constrained flock animations. Copyright © 2008 John Wiley & Sons, Ltd.
Jiayi Xu 0002, Xiaogang Jin 0001, Yizhou Yu, Tian Shen, Mingdong Zhou
Comput. Animat. Virtual Worlds1
2007 Interactive control of real-time crowd navigation in virtual environment
abstract
Interactive control is one of the key issues when simulating crowd navigation in virtual environment. In this paper, we propose a simple but practical method for authoring crowd scenes in an effective and intuitive way. Radial Basis Functions (RBF) based vector field is employed as the governing tool to drive the motion flow. With this basic mathematical tool, users can easily control the motions of crowd by simply sketching velocities on a few points in the scene. Our approach is fast enough to allow on-the-fly modification of the vector field. Besides, the behavior of an individual in a crowd can be interactively adjusted by changing the ratio between its autonomous and governed movements.
Xiaogang Jin 0001, Charlie C. L. Wang, Shengsheng Huang, Jiayi Xu 0002
VRST4