Kangneng Zhou

dblp:257/7652 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-9102-9527ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng
Int. J. Comput. Vis.1
2025 VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior
abstract
Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, facial motion, head pose generation, and video quality. However, no model has yet led or tied on all these metrics due to the one-to-many mapping between audio and motion. In this paper, we propose VividTalk, a two-stage generic framework that supports generating high-visual quality talking head videos with all the above properties. Specifically, in the first stage, we map the audio to mesh by learning two motions, including non-rigid facial motion and rigid head motion. For facial motion, both blendshape and vertex are adopted as the intermediate representation to maximize the representation ability of the model. For head motion, a novel learnable head pose codebook with a two-phase training mechanism is proposed. In the second stage, we proposed a dual branch motion-vae and a generator to transform the meshes into dense motion and synthesize high-quality video frame-by-frame. Extensive experiments show that the proposed VividTalk can generate high-visual quality talking head videos with lip-sync and realistic enhanced by a large margin, and outperforms previous state-of-the-art works in objective and subjective comparisons. The code will be publicly released upon publication.
Xusen Sun, Longhao Zhang, Hao Zhu 0004, Peng Zhang 0080, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, Xun Cao
3DV7
2025 3DFaceController: Region-Controllable Face Synthesis via Decomposed and Recomposed Neural Radiance Fields
Kangneng Zhou, Yaxing Wang, Shuang Song 0005, Jie Zhang 0090, Ping Li 0016
CVM (2)1
2025 3DCMM: 3D Comprehensive Morphable Models With UV-UNet for Accurate Head Creation
abstract
In recent studies of 3D shape modelling and reconstruction, the focus has primarily been on the 3D face region. However, accurately creating the entire 3D head opens up a wide range of applications, including headwear design, cranial diagnosis, and avatar design. Therefore, we present our newly developed method of constructing 3D comprehensive morphable models (3DCMM) specifically tailored for human heads, along with a novel 3DCMM-based stepwise pipeline for creating accurate full 3D heads. Within our 3DCMM framework, we constructed a powerful 3D morphable face model with UV-UNet to generate the 3D face and predict the 3D scalp, resulting in a complete representation of the head. Additionally, our 3DCMM-based self-learning approach incorporates novel facial boundary-aware and structure-aware losses for highly accurate overall reconstructions of the entire facial region. Experimental evaluations demonstrate that our 3DCMM exhibits superior face representation power and achieves higher head prediction accuracy than existing models. Consequently, our 3DCMM-based 3D head creation method from a single image demonstrates outstanding performance capability on both face and head benchmarks.
Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Tong-Yee Lee, Ping Li 0016
IEEE Trans. Multim.2
2024 Multi-view X-ray Image Synthesis with Multiple Domain Disentanglement from CT Scans
abstract
X-ray images play a vital role in the intraoperative processes due to their high resolution and fast imaging speed and greatly promote the subsequent segmentation, registration and reconstruction. However, over-dosed X-rays superimpose potential risks to human health to some extent. Data-driven algorithms from volume scans to X-ray images are restricted by the scarcity of paired X-ray and volume data. Existing methods are mainly realized by modelling the whole X-ray imaging procedure. In this study, we propose a learning-based approach termed CT2X-GAN to synthesize the X-ray images in an end-to-end manner using the content and style disentanglement from three different image domains. Our method decouples the anatomical structure information from CT scans and style information from unpaired real X-ray images/ digital reconstructed radiography (DRR) images via a series of decoupling encoders. Additionally, we introduce a novel consistency regularization term to improve the stylistic resemblance between synthesized X-ray images and real X-ray images. Meanwhile, we also impose a supervised process by computing the similarity of computed real DRR and synthesized DRR images. We further develop a pose attention module to fully strengthen the comprehensive information in the decoupled content code from CT scans, facilitating high-quality multi-view image synthesis in the lower 2D space. Extensive experiments were conducted on the publicly available CTSpine1K dataset and achieved 97.8350, 0.0842 and 3.0938 in terms of FID, KID and defined user-scored X-ray similarity, respectively. In comparison with 3D-aware methods (π-GAN, EG3D), CT2X-GAN is superior in improving the synthesis quality and realistic to the real X-ray images.
Lixing Tan, Shuang Song 0005, Kangneng Zhou, Chengbo Duan, Huayang Ren, Wei Zhang 0373, Ruoxiu Xiao
ACM Multimedia3
2024 MeshWGAN: Mesh-to-Mesh Wasserstein GAN With Multi-Task Gradient Penalty for 3D Facial Geometric Age Transformation
abstract
As the metaverse develops rapidly, 3D facial age transformation is attracting increasing attention, which may bring many potential benefits to a wide variety of users, e.g., 3D aging figures creation, 3D facial data augmentation and editing. Compared with 2D methods, 3D face aging is an underexplored problem. To fill this gap, we propose a new mesh-to-mesh Wasserstein generative adversarial network (MeshWGAN) with a multi-task gradient penalty to model a continuous bi-directional 3D facial geometric aging process. To the best of our knowledge, this is the first architecture to achieve 3D facial geometric age transformation via real 3D scans. As previous image-to-image translation methods cannot be directly applied to the 3D facial mesh, which is totally different from 2D images, we built a mesh encoder, decoder, and multi-task discriminator to facilitate mesh-to-mesh transformations. To mitigate the lack of 3D datasets containing children's faces, we collected scans from 765 subjects aged 5-17 in combination with existing 3D face databases, which provided a large training dataset. Experiments have shown that our architecture can predict 3D facial aging geometries with better identity preservation and age closeness compared to 3D trivial baselines. We also demonstrated the advantages of our approach via various 3D face-related graphics applications.
Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Tong-Yee Lee, Ping Li 0016
IEEE Trans. Vis. Comput. Graph.2
2023 KSCB: a novel unsupervised method for text sentiment analysis
Weili Jiang, Kangneng Zhou, Chenchen Xiong, Guodong Du 0002, Chubin Ou, Junpeng Zhang 0001
Appl. Intell.2
2023 Generative Consistency for Semi-Supervised Cerebrovascular Segmentation From TOF-MRA
abstract
Cerebrovascular segmentation from Time-of-flight magnetic resonance angiography (TOF-MRA) is a critical step in computer-aided diagnosis. In recent years, deep learning models have proved its powerful feature extraction for cerebrovascular segmentation. However, they require many labeled datasets to implement effective driving, which are expensive and professional. In this paper, we propose a generative consistency for semi-supervised (GCS) model. Considering the rich information contained in the feature map, the GCS model utilizes the generation results to constrain the segmentation model. The generated data comes from labeled data, unlabeled data, and unlabeled data after perturbation, respectively. The GCS model also calculates the consistency of the perturbed data to improve the feature mining ability. Subsequently, we propose a new model as the backbone of the GSC model. It transfers TOF-MRA into graph space and establishes correlation using Transformer. We demonstrated the effectiveness of the proposed model on TOF-MRA representations, and tested the GCS model with state-of-the-art semi-supervised methods using the proposed model as backbone. The experiments prove the important role of the GCS model in cerebrovascular segmentation. Code is available at https://github.com/MontaEllis/SSL-For-Medical-Segmentation.
Cheng Chen 0024, Kangneng Zhou, Ruoxiu Xiao
IEEE Trans. Medical Imaging2
2022 SD-GAN: Semantic Decomposition for Face Image Synthesis with Discrete Attribute
abstract
Manipulating latent code in generative adversarial networks (GANs) for facial image synthesis mainly focuses on continuous attribute synthesis (e.g., age, pose and emotion), while discrete attribute synthesis (like face mask and eyeglasses) receives less attention. Directly applying existing works to facial discrete attributes may cause inaccurate results. In this work, we propose an innovative framework to tackle challenging facial discrete attribute synthesis via semantic decomposing, dubbed SD-GAN. To be concrete, we explicitly decompose the discrete attribute representation into two components, i.e. the semantic prior basis and offset latent representation. The semantic prior basis shows an initializing direction for manipulating face representation in the latent space. The offset latent presentation obtained by 3D-aware semantic fusion network is proposed to adjust prior basis. In addition, the fusion network integrates 3D embedding for better identity preservation and discrete attribute synthesis. The combination of prior basis and offset latent representation enable our method to synthesize photo-realistic face images with discrete attributes. Notably, we construct a large and valuable dataset MEGN (Face Mask and Eyeglasses images crawled from Google and Naver) for completing the lack of discrete attributes in the existing dataset. Extensive qualitative and quantitative experiments demonstrate the state-of-the-art performance of our method. Our code is available at an anonymous website: https://github.com/MontaEllis/SD-GAN.
Kangneng Zhou, Xiaobin Zhu 0001, Daiheng Gao, Kai Lee, Xinjie Li 0002, Xu-Cheng Yin
ACM Multimedia1
2022 Customize My Helmet: A Novel Algorithmic Approach Based on 3D Head Prediction
Jie Zhang 0090, Yan Luximon, Parth B. Shah, Kangneng Zhou, Ping Li 0016
Comput. Aided Des.4
2022 3D-guided facial shape clustering and analysis
Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Ping Li 0016, Hassan Iftikhar
Multim. Tools Appl.2
2021 An Effective Deep Neural Network for Lung Lesions Segmentation From COVID-19 CT Images
abstract
Automatic segmentation of lung lesions from COVID-19 computed tomography (CT) images can help to establish a quantitative model for diagnosis and treatment. For this reason, this article provides a new segmentation method to meet the needs of CT images processing under COVID-19 epidemic. The main steps are as follows: First, the proposed region of interest extraction implements patch mechanism strategy to satisfy the applicability of 3-D network and remove irrelevant background. Second, 3-D network is established to extract spatial features, where 3-D attention model promotes network to enhance target area. Then, to improve the convergence of network, a combination loss function is introduced to lead gradient optimization and training direction. Finally, data augmentation and conditional random field are applied to realize data resampling and binary segmentation. This method was assessed with some comparative experiment. By comparison, the proposed method reached the highest performance. Therefore, it has potential clinical applications.
Cheng Chen 0024, Kangneng Zhou, Muxi Zha, Xiangyan Qu, Xiaoyu Guo 0004, Ruoxiu Xiao
IEEE Trans. Ind. Informatics2