Jichao Zhang

dblp:84/8496 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 High-Fidelity 3D Facial Avatar Synthesis With Controllable Fine-Grained Expressions
abstract
Facial expression editing methods can be mainly categorized into two types based on their architectures: 2D-based and 3D-based methods. The former lacks 3D face modeling capabilities, making it difficult to edit 3D factors effectively. The latter has demonstrated superior performance in generating high-quality and view-consistent renderings using single-view 2D face images. Although these methods have successfully used animatable models to control facial expressions, they still have limitations in achieving precise control over fine-grained expressions. To address this issue, in this paper, we propose a novel approach by simultaneously refining both the latent code of a pretrained 3D-Aware GAN model for texture editing and the expression code of the driven 3DMM model for mesh editing. Specifically, we introduce a Dual Mappers module, comprising Texture Mapper and Emotion Mapper, to learn the transformations of the given latent code for textures and the expression code for meshes, respectively. To optimize the Dual Mappers, we propose a Text-Guided Optimization method, leveraging a CLIP-based objective function with expression text prompts as targets, while integrating a SubSpace Projection mechanism to project the text embedding to the expression subspace such that we can have more precise control over fine-grained expressions. Extensive experiments and comparative analyses demonstrate the effectiveness and superiority of our proposed method.
Yikang He, Jichao Zhang, Wei Wang 0108, Nicu Sebe, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 DualSpec: Text-to-Spatial-Audio Generation via Dual-Spectrogram Guided Diffusion Model
abstract
Text-to-audio (TTA), which generates audio signals from textual descriptions, has received huge attention in recent years. However, recent works focused on text to monaural audio only. As we know, spatial audio provides more immersive auditory experience than monaural audio, e.g. in virtual reality. To address this issue, we propose a text-to-spatial-audio (TTSA) generation framework named DualSpec. Specifically, it first trains variational autoencoders (VAEs) for extracting the latent acoustic representations from sound event audio. Then, given text that describes sound events and event directions, the proposed method uses the encoder of a pretrained large language model to transform the text into text features. Finally, it trains a diffusion model from the latent acoustic representations and text features for the spatial audio generation. In the inference stage, only the text description is needed to generate spatial audio. Particularly, to improve the synthesis quality and azimuth accuracy of the spatial sound events simultaneously, we propose to use two kinds of acoustic features. One is the Mel spectrograms which is good for improving the synthesis quality, and the other is the short-time Fourier transform spectrograms which is good at improving the azimuth accuracy. We provide a pipeline of constructing spatial audio dataset with text prompts, for the training of the VAEs and diffusion model. We also introduce new spatial-aware evaluation metrics to quantify the azimuth errors of the generated spatial audio recordings. Experimental results demonstrate that the proposed method can generate spatial audio with high directional and event consistency.
Lei Zhao 0031, Sizhou Chen, Linfeng Feng, Jichao Zhang, Xiao-Lei Zhang 0001, Xuelong Li 0001
IEEE Trans. Multim.4
2026 Speech-Driven 3D Facial Animation with Natural Head Movements
abstract
Speech-driven 3D facial animation has been studied for a long time, yet many challenges still prevent achieving truly natural results. One major issue is that current methods often overlook head motion during speech. To address this, we conducted a preliminary investigation and identified several key obstacles: the scarcity of 3D facial animation datasets with head motion and the limitations of regular sequence prediction models (e.g., Transformer, LSTM, and GRU), which are designed for discrete dynamic sequences and do not align well with continuous 3D head motion sequences. To solve these issues, we propose a head motion prediction module. This module uses audio and the initial motion state of the mesh to predict head motion. By employing ordinary differential equations (ODE) to model continuous dynamic sequences, it predicts head movements that closely resemble real head motion, making the animation more realistic and natural. Additionally, recognizing that the head motion generation given audio is not a one-to-one mapping problem, we introduce a noise module to help the head motion prediction module generate varied head motions given the same audio input. We also observed that previous facial animation methods primarily focus on generating vertices for the mouth region but use a single model to generate the entire face. This approach wastes some of the model’s fitting capacity on other regions. To solve this problem, we propose a cascaded mesh generation module that uses two modules to separately generate the vertex of mouth region and other facial regions. Extensive experiments and a perceptual user study show that our approach outperforms existing methods and produces relatively natural head motion.
Kuiyuan Sun, Jichao Zhang, Wei Wang 0108, Yao Zhao 0001, Nicu Sebe
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Stable-Hair V2: Real-World Hair Transfer via Multiple-View Diffusion Model
abstract
While diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outputs-crucial for real-world applications such as digital humans and virtual avatars-remains underexplored. In this paper, we propose Stable-Hair v2, a novel diffusion-based multi-view hair transfer framework. To the best of our knowledge, this is the first work to leverage multiple-view diffusion models for robust, high-fidelity, and view-consistent hair transfer across multiple perspectives. We introduce a comprehensive multi-view training data generation pipeline to generate high-quality triplet data, including bald images, reference hairstyles, and view-aligned source-bald pairs. Our multi-view hair transfer model integrates polar-azimuth embeddings for pose conditioning and temporal attention layers to ensure smooth transitions between views. To optimize this model, we design a novel multi-stage training strategy consisting of Pose-Controllable Latent IdentityNet training, Hair Extractor training, and Temporal Attention training. Extensive experiments demonstrate that our method accurately transfers detailed and realistic hairstyles to source subjects while achieving seamless and consistent results across views, significantly outperforming existing methods and establishing a new benchmark in multi-view hair transfer.
Kuiyuan Sun, Jichao Zhang, Wei Wang 0108, Nicu Sebe, Yao Zhao 0001
IEEE Trans. Vis. Comput. Graph.3
2025 Stable-Hair: Real-World Hair Transfer via Diffusion Model
abstract
Current hair transfer methods struggle to handle diverse and intricate hairstyles, limiting their applicability in real-world scenarios. In this paper, we propose a novel diffusion-based hair transfer framework, named Stable-Hair, which robustly transfers a wide range of real-world hairstyles to user-provided faces for virtual hair try-on. To achieve this goal, our Stable-Hair framework is designed as a two-stage pipeline. In the first stage, we train a Bald Converter alongside stable diffusion to remove hair from the user-provided face images, resulting in bald images. In the second stage, we specifically designed a Hair Extractor and a Latent IdentityNet to transfer the target hairstyle with highly detailed and high-fidelity to the bald image. The Hair Extractor is trained to encode reference images with the desired hairstyles, while the Latent IdentityNet ensures consistency in identity and background. To minimize color deviations between source images and transfer results, we introduce a novel Latent ControlNet architecture, which functions as both the Bald Converter and Latent IdentityNet. After training on our curated triplet dataset, our method accurately transfers highly detailed and high-fidelity hairstyles to the source images. Extensive experiments demonstrate that our approach achieves state-of-the-art performance compared to existing hair transfer methods.
Yuxuan Zhang 0001, Yiren Song, Jichao Zhang, Hao Tang 0005
AAAI4
2025 Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
abstract
The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, the compression process of LDM often results in the deterioration of details, particularly in sensitive areas such as facial features and clothing textures. In this paper, we propose a Multi-focal Conditioned Latent Diffusion (MCLD) method to address these limitations by conditioning the model on disentangled, pose-invariant features from these sensitive regions. Our approach utilizes a multi-focal condition aggregation module, which effectively integrates facial identity and texture-specific information, enhancing the model’s ability to produce appearance realistic and identity-consistent images. Our method demonstrates consistent identity and appearance generation on the Deep-Fashion dataset and enables flexible person image editing due to its generation consistency. The code is available at https://github.com/jqliu09/mcld.
Jichao Zhang, Paolo Rota, Nicu Sebe
CVPR2
2025 Novel SMC for Discrete Interval Type-2 Fuzzy Semi-Markovian Switching Models With Incomplete Semi-Markovian Kernel
abstract
This work studies the novel sliding mode control (SMC) of discrete nonlinear stochastic switching models under semi-Markovian parameter and incomplete semi-Markovian kernel (SMK). The characteristic of nonlinear system is described by an interval type-2 fuzzy (IT2F) model that can be recognized as a collection of several type-1 fuzzy models. The uncertainties in system parameters is efficiently captured using the lower and upper grades of membership. Based on the mode-dependent Lyapunov function and incomplete SMK, sufficient conditions are proposed to ensure the stability of sliding dynamics. Moreover, an IT2F SMC law based on learning strategy is developed such that the state signals are guided onto the predetermined sliding region and the mode switchings-induced chattering is effectively reduced. Finally, the IT2F SMC strategy is validated through the simulation of truck-trailer model.
Wenhai Qi, Jichao Zhang, Guangdeng Zong, Shun-Feng Su, Jinde Cao, Ruey-Huei Yeh
IEEE Trans. Cybern.2
2024 UVMap-ID: A Controllable and Personalized UV Map Generative Model
abstract
Recently, diffusion models have made significant strides in synthesizing realistic 2D human images based on provided text prompts. Building upon this, researchers have extended 2D text-to-image diffusion models into the 3D domain for generating human textures (UV Maps). However, some important problems about UV Map Generative models are still not solved, i.e., how to generate personalized texture maps for any given face image, and how to define and evaluate the quality of these generated texture maps. To solve the above problems, we introduce a novel method, UVMap-ID, which is a controllable and personalized UV Map generative model. Unlike traditional large-scale training methods in 2D, we propose to fine-tune a pre-trained text-to-image diffusion model which is integrated with a face fusion module for achieving ID-driven customized generation. To support the finetuning strategy, we introduce a small-scale attribute-balanced training dataset, including high-quality textures with labeled text and Face ID. Additionally, we introduce some metrics to evaluate the multiple aspects of the textures. Finally, both quantitative and qualitative analyses demonstrate the effectiveness of our method in controllable and personalized UV Map generation.
Weijie Wang 0002, Jichao Zhang, Chang Liu 0030, Xia Li 0005, Xingqian Xu, Humphrey Shi, Nicu Sebe, Bruno Lepri
ACM Multimedia2
2024 Fuzzy SMC for Discrete Nonlinear Singularly Perturbed Models With Semi-Markovian Switching Parameters
abstract
This paper investigates the sliding mode control (SMC) issue for discrete nonlinear semi-Markovian switching singularly perturbed models with incomplete semi-Markovian kernels (SMKs). A T-S fuzzy model is adopted to characterize the underlying model with singularly perturbed parameters, which is more widely used in nonlinear systems. Because the statistical characteristic of a semi-Markovian kernel is laborious to be obtained in practical engineering, the information of semi- Markov kernel is recognized to be incomplete. The main novelty lies in constructing a suitable fuzzy SMC scheme to realize the attainment of the quasi-sliding mode, overcoming the difficulty caused by incomplete SMK. Based on the Lyapunov function related to the fuzzy rule and the elapsed time of current mode, the stability criterion is constructed for the corresponding model. An appropriate SMC approach is then developed to achieve the attainment of a quasi-sliding mode. Finally, the theoretical results are verified by an improved tunnel diode circuit model.
Wenhai Qi, Jichao Zhang, Ju H. Park 0001, Hak-Keung Lam, Jun Cheng 0004
IEEE Trans. Fuzzy Syst.2
2023 Householder Projector for Unsupervised Latent Semantics Discovery
abstract
Generative Adversarial Networks (GANs), especially the recent style-based generators (StyleGANs), have versatile semantics in the structured latent space. Latent semantics discovery methods emerge to move around the latent code such that only one factor varies during the traversal. Recently, an unsupervised method proposed a promising direction to directly use the eigenvectors of the projection matrix that maps latent codes to features as the interpretable directions. However, one overlooked fact is that the projection matrix is non-orthogonal and the number of eigenvectors is too large. The non-orthogonality would entangle semantic attributes in the top few eigenvectors, and the large dimensionality might result in meaningless variations among the directions even if the matrix is orthogonal. To avoid these issues, we propose Householder Projector, a flexible and general low-rank orthogonal matrix representation based on Householder transformations, to parameterize the projection matrix. The orthogonality guarantees that the eigenvectors correspond to disentangled interpretable semantics, while the low-rank property encourages that each identified direction has meaningful variations. We integrate our projector into pre-trained StyleGAN2/StyleGAN3 and evaluate the models on several benchmarks. Within only 1% of the original training steps for fine-tuning, our projector helps StyleGANs to discover more disentangled and precise semantic attributes without sacrificing image fidelity. Code is publicly available via https://github.com/KingJamesSong/HouseholderGAN.
Yue Song 0002, Jichao Zhang, Nicu Sebe, Wei Wang 0108
ICCV2
2023 SMC for Discrete Fuzzy Semi-Markov Jump Models With Partly Known Semi-Markov Kernel
abstract
This article investigates the sliding mode control (SMC) for discrete-time nonlinear semi-Markov jump models with a partly known semi-Markov kernel (SMK). The nonlinear system is characterized by the Takagi–Sugeno (T–S) fuzzy model, where the membership functions for fuzzy rules are designed to be related to the system mode. In view of the fact that the statistical characteristic of the SMK is difficult to fully obtain in practical engineering, the SMK is recognized to be partly known with less conservativeness than both semi-Markov jump models with completely known SMK and Markov jump models with partly known transition probabilities. On the basis of classical Lyapunov stability and fuzzy-model-based approach, novel convex mean-square stability is proposed for the underlying system by eliminating the nonlinear coupling terms with the aid of additional matrix variables. Afterward, a fuzzy SMC law strategy is constructed to guarantee the reachability of the discrete quasi-sliding mode. Finally, a robot arm model is simulated to verify the proposed fuzzy SMC strategy.
Wenhai Qi, Jichao Zhang, Ju H. Park 0001, Zhengguang Wu, Huaicheng Yan 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2022 3D-Aware Semantic-Guided Generative Model for Human Synthesis
Jichao Zhang, Enver Sangineto, Hao Tang 0005, Aliaksandr Siarohin, Zhun Zhong, Nicu Sebe, Wei Wang 0108
ECCV (15)1
2022 Unsupervised High-Resolution Portrait Gaze Correction and Animation
abstract
This paper proposes a gaze correction and animation method for high-resolution, unconstrained portrait images, which can be trained without the gaze angle and the head pose annotations. Common gaze-correction methods usually require annotating training data with precise gaze, and head pose information. Solving this problem using an unsupervised method remains an open problem, especially for high-resolution face images in the wild, which are not easy to annotate with gaze and head pose labels. To address this issue, we first create two new portrait datasets: CelebGaze ( 256 ×256 ) and high-resolution CelebHQGaze ( 512 ×512 ). Second, we formulate the gaze correction task as an image inpainting problem, addressed using a Gaze Correction Module (GCM) and a Gaze Animation Module (GAM). Moreover, we propose an unsupervised training strategy, i.e., Synthesis-As-Training, to learn the correlation between the eye region features and the gaze angle. As a result, we can use the learned latent space for gaze animation with semantic interpolation in this space. Moreover, to alleviate both the memory and the computational costs in the training and the inference stage, we propose a Coarse-to-Fine Module (CFM) integrated with GCM and GAM. Extensive experiments validate the effectiveness of our method for both the gaze correction and the gaze animation tasks in both low and high-resolution face datasets in the wild and demonstrate the superiority of our method with respect to the state of the art.
Jichao Zhang, Hao Tang 0005, Enver Sangineto, Peng Wu 0014, Yan Yan 0002, Nicu Sebe, Wei Wang 0108
IEEE Trans. Image Process.1
2021 Coarse-to-Fine Gaze Redirection with Numerical and Pictorial Guidance
abstract
Gaze redirection aims at manipulating the gaze of a given face image with respect to a desired direction (i.e., a reference angle) and it can be applied to many real life scenarios, such as video-conferencing or taking group photos. However, previous work on this topic mainly suffers of two limitations: (1) Low-quality image generation and (2) Low redirection precision. In this paper, we propose to alleviate these problems by means of a novel gaze redirection framework which exploits both a numerical and a pictorial direction guidance, jointly with a coarse-to-fine learning strategy. Specifically, the coarse branch learns the spatial transformation which warps input image according to desired gaze. On the other hand, the fine-grained branch consists of a generator network with conditional residual image learning and a multi-task discriminator. This second branch reduces the gap between the previously warped image and the ground-truth image and recovers finer texture details. Moreover, we propose a numerical and pictorial guidance module (NPG) which uses a pictorial gazemap description and numerical angles as an extra guide to further improve the precision of gaze redirection. Extensive experiments on a benchmark dataset show that the proposed method outperforms the state-of-the-art approaches in terms of both image quality and redirection precision. The code is available at https://github.com/jingjingchen777/CFGR
Jichao Zhang, Enver Sangineto, Tao Chen 0003, Jiayuan Fan 0001, Nicu Sebe
WACV2
2020 Dual In-painting Model for Unsupervised Gaze Correction and Animation in the Wild
abstract
We address the problem of unsupervised gaze correction in the wild, presenting a solution that works without the need of precise annotations of the gaze angle and the head pose. We created a new dataset called CelebAGaze consisting of two domains X, Y, where the eyes are either staring at the camera or somewhere else. Our method consists of three novel modules: the Gaze Correction module(GCM), the Gaze Animation module(GAM), and the Pretrained Autoencoder module (PAM). Specifically, GCM and GAM separately train a dual in-painting network using data from the domain X for gaze correction and data from the domain Y for gaze animation. Additionally, a Synthesis-As-Training method is proposed when training GAM to encourage the features encoded from the eye region to be correlated with the angle information, resulting in gaze animation achieved by interpolation in the latent space. To further preserve the identity information e.g., eye shape, iris color, we propose the PAM with an Autoencoder, which is based on Self-Supervised mirror learning where the bottleneck features are angle-invariant and which works as an extra input to the dual in-painting models. Extensive experiments validate the effectiveness of the proposed method for gaze correction and gaze animation in the wild and demonstrate the superiority of our approach in producing more compelling results than state-of-the-art baselines. Our code, the pretrained models and supplementary results are available at:https://github.com/zhangqianhui/GazeAnimation.
Jichao Zhang, Hao Tang 0005, Wei Wang 0108, Yan Yan 0002, Enver Sangineto, Nicu Sebe
ACM Multimedia1
2018 Sparsely Grouped Multi-Task Generative Adversarial Networks for Facial Attribute Manipulation
abstract
Recently, Image-to-Image Translation (IIT) has achieved great progress in image style transfer and semantic context manipulation for images. However, existing approaches require exhaustively labelling training data, which is labor demanding, difficult to scale up, and hard to adapt to a new domain. To overcome such a key limitation, we propose Sparsely Grouped Generative Adversarial Networks (SG-GAN) as a novel approach that can translate images in sparsely grouped datasets where only a few train samples are labelled. Using a one-input multi-output architecture, SG-GAN is well-suited for tackling multi-task learning and sparsely grouped learning tasks. The new model is able to translate images among multiple groups using only a single trained model. To experimentally validate the advantages of the new model, we apply the proposed method to tackle a series of attribute manipulation tasks for facial images as a case study. Experimental results show that SG-GAN can achieve comparable results with state-of-the-art methods on adequately labelled datasets while attaining a superior image translation quality on sparsely grouped datasets~\footnoteCode is available at https://github.com/zhangqianhui/SGGAN-tensorflow..
Jichao Zhang, Yezhi Shu, Songhua Xu, Gongze Cao, Fan Zhong 0001, Meng Liu 0006, Xueying Qin
ACM Multimedia1
2017 ST-GAN: Unsupervised Facial Image Semantic Transformation Using Generative Adversarial Networks
abstract
Image semantic transformation aims to convert one image into another image with different semantic features (e.g., face pose, hairstyle). The previous methods, which learn the mapping function from one image domain to the other, require supervised information directly or indirectly. In this paper, we propose an unsupervised image semantic transformation method called semantic transformation generative adversarial networks (ST-GAN), and experimentally verify it on face dataset. We further improve ST-GAN with the Wasserstein distance to generate more realistic images and propose a method called local mutual information maximization to obtain a more explicit semantic transformation. ST-GAN has the ability to map the image semantic features into the latent vector and then perform transformation by controlling the latent vector.
Jichao Zhang, Fan Zhong 0001, Gongze Cao, Xueying Qin
ACML1