Zhe Kong

dblp:154/1597 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
abstract
Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM
Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen
CVPR4
2025 Scalable Dual Fingerprinting for Hierarchical Attribution of Text-to-Image Models
Jianwei Fei, Yunshu Dai, Peipeng Yu, Zhe Kong, Zhihua Xia
ICCV4
2025 MOERL: When Mixture-Of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration
Tao Wang 0052, Peiwen Xia, Peng-Tao Jiang, Zhe Kong, Kaihao Zhang, Tong Lu 0002, Wenhan Luo
ICCV5
2025 FineMotion: A Dataset and Benchmark with Both Spatial and Temporal Annotation for Fine-Grained Motion Generation and Editing
Bizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen
ICCV4
2025 Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
abstract
Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus on single human animation and struggle with multi-stream audio inputs, facing incorrect binding problems between audio and persons. Additionally, they exhibit limitations in instruction-following capabilities. To solve this problem, in this paper, we propose a novel task: Multi-Person Conversational Video Generation, and introduce a new framework, MultiTalk, to address the challenges during multi-person generation. Specifically, for audio injection, we investigate several schemes and propose the Label Rotary Position Embedding (L-RoPE) method to resolve the audio and person binding problem. Furthermore, during training, we observe that partial parameter training and multi-task training are crucial for preserving the instruction-following ability of the base model. MultiTalk achieves superior performance compared to other methods on several datasets, including talking head, talking body, and multi-person datasets, demonstrating the powerful generation capabilities of our approach.
Zhe Kong, Yong Zhang 0034, Zhuoliang Kang, Xiaoming Wei, Guanying Chen, Wenhan Luo
NeurIPS1
2024 OMG: Occlusion-Friendly Personalized Multi-concept Generation in Diffusion Models
Zhe Kong, Yong Zhang 0034, Tianyu Yang 0003, Tao Wang 0052, Kaihao Zhang, Bizhu Wu, Guanying Chen, Wei Liu 0005, Wenhan Luo
ECCV (31)1
2024 Enhancing Document-Level Event Extraction via Structure-Aware Heterogeneous Graph with Multi-Granularity Subsentences
abstract
Document-level Event Extraction aims to identify events from an entire article. It is quite a challenging task because event arguments scatter across several sentences and multiple events in a document may have influence on each other. Previous methods, however, did not take advantage of document structures that have been proved to be effective for sentence-level event extraction. In this work, we propose a structure-aware heterogeneous graph with subsentences for document-level event extraction. Firstly, we build a syntactic graph to capture long-range dependencies between cross-sentence event arguments. Then, multi-granularity sub-sentences are added into the graph to acquire fine-grained understanding. Finally, a global memory stores extracted events so that interactions among multiple events can be captured. Extensive experiments demonstrate that our model outperforms the state of the art models on a widely used large-scale document-level event extraction dataset.
Yuhan Liu 0012, Neng Gao, Zhe Kong
ICASSP4
2024 Enhancing Generative Generalized Zero Shot Learning via Multi-Space Constraints and Adaptive Integration
Zhe Kong, Neng Gao, Yuhan Liu 0012
MMM (1)1
2024 Dual Teacher Knowledge Distillation With Domain Alignment for Face Anti-Spoofing
abstract
Face recognition systems have raised concerns due to their vulnerability to different presentation attacks, and system security has become an increasingly critical concern. Although many face anti-spoofing (FAS) methods perform well in intra-dataset scenarios, their generalization remains a challenge. To address this issue, some methods adopt domain adversarial training (DAT) to extract domain-invariant features. Differently, in this paper, we propose a domain adversarial attack (DAA) method by adding perturbations to the input images, which makes them indistinguishable across domains and enables domain alignment. Moreover, since models trained on limited data and types of attacks cannot generalize well to unknown attacks, we propose a dual perceptual and generative knowledge distillation framework for face anti-spoofing that utilizes pre-trained face-related models containing rich face priors. Specifically, we adopt two different face-related models as teachers to transfer knowledge to the target student model. The pre-trained teacher models are not from the task of face anti-spoofing but from perceptual and generative tasks, respectively, which implicitly augment the data. By combining both DAA and dual-teacher knowledge distillation, we develop a dual teacher knowledge distillation with domain alignment framework (DTDA) for face anti-spoofing. The advantage of our proposed method has been verified through extensive ablation studies and comparison with state-of-the-art methods on public datasets across multiple protocols.
Zhe Kong, Wentian Zhang, Tao Wang 0052, Kaihao Zhang, Yuexiang Li, Xiaoying Tang 0001, Wenhan Luo
IEEE Trans. Circuits Syst. Video Technol.1
2024 Taming Self-Supervised Learning for Presentation Attack Detection: De-Folding and De-Mixing
abstract
Biometric systems are vulnerable to presentation attacks (PAs) performed using various PA instruments (PAIs). Even though there are numerous PA detection (PAD) techniques based on both deep learning and hand-crafted features, the generalization of PAD for unknown PAI is still a challenging problem. In this work, we empirically prove that the initialization of the PAD model is a crucial factor for generalization, which is rarely discussed in the community. Based on such observation, we proposed a self-supervised learning-based method, denoted as DF-DM. Specifically, DF-DM is based on a global-local view coupled with de-folding and de-mixing to derive the task-specific representation for PAD. During de-folding, the proposed technique will learn region-specific features to represent samples in a local pattern by explicitly minimizing the generative loss. While de-mixing drives detectors to obtain the instance-specific features with global information for more comprehensive representation by minimizing the interpolation-based consistency. Extensive experimental results show that the proposed method can achieve significant improvements in terms of both face and fingerprint PAD in more complicated and hybrid datasets when compared with the state-of-the-art methods. When training in CASIA-FASD and Idiap Replay-Attack, the proposed method can achieve an 18.60% equal error rate (EER) in OULU-NPU and MSU-MFSD, exceeding the baseline performance by 9.54%. The source code of the proposed technique is available at https://github.com/kongzhecn/dfdm.
Zhe Kong, Wentian Zhang, Feng Liu 0013, Wenhan Luo, LinLin Shen, Ramachandra Raghavendra
IEEE Trans. Neural Networks Learn. Syst.1
2022 A Noise-Aware Framework for Blind Image Super-Resolution
abstract
The real-world image degradation in the super-resolution task is recently considered as a combination of Gaussian blur, down-sampling, and additional white Gaussian noise. To han-dle this degradation, previous methods estimate the Gaussian blur kernel or model the degradation based on a randomly selected image patch. However, these methods cannot han-dle degradations with high-level noise well as they ignore the spatial variability or even the existence of noise. Moreover, using image denoising networks to preprocess low-resolution images also fails due to the loss of important high-frequency information. In this paper, we propose a framework called EASE to flexibly handle real-world degradations. Specifi-cally, we develop a lightweight module to erase noise and blur simultaneously by learning from an image denoising and an image restoration network, which adapts to existing net-works that focus on handling bicubic down-sampling. Exten-sive experiments prove the superiority of our method, espe-cially when handling degradations with high-level noise.
Guanqun Liu 0002, Xin Wang 0086, Lei Wang 0135, Daren Zha, Lin Zhao 0006, Zhe Kong, Peng Qi 0005
ICME6
2022 Searching Models with Nested Attention for Blind Super-Resolution
abstract
Blind super-resolution task aims to restore low-resolution im“ages with unknown degradations to high-resolution counter-parts. Existing methods rely on degradations estimation to re-construct high-resolution images. However, they need human involvement to obtain the best results as they treat unknown types of degradations as known conditions and manually select corresponding trained models. Moreover, they cannot fully use estimated degradations and generate blurry artifacts as they ignore that the impact of degradations on images is re-lated to images contents. In this paper, we propose HIS-NEST which contains an automatic search strategy HIS and a net-work structure NEST. Specifically, to bypass manual partici-pation, HIS automatically selects the clearest image by esti-mating the qualities of generated images. Furthermore, NEST protects the connection between degradations and images by using no loss functions to limit the degradations estimation and analyzing degradations from the perspective of channel and space. Extensive experiments show that our method out-performs state-of-the-art methods.
Guanqun Liu 0002, Xin Wang 0086, Lei Wang 0135, Daren Zha, Lin Zhao 0006, Zhe Kong, Peng Qi 0005
ICME6
2022 Multi-level Fusion of Multi-modal Semantic Embeddings for Zero Shot Learning
abstract
Zero shot learning aims to recognize objects whose instances may not be covered by the training data. To generalize knowledge from seen classes to the novel ones, semantic space is built to embed knowledge from various views into multi-modal semantic embeddings. Existing semantic embeddings neglect the relationships between classes which are essential to transfer knowledge between classes. Moreover, existing zero shot learning models ignore the complementarity between semantic embeddings from different modalities. To tackle these problems, in this work, we resort to graph theory to explicitly model the interdependence between classes and then obtain new modal semantic embeddings. Furthermore, we pioneer to propose a multi-level fusion model to effectively combine knowledge encoded in multi-modal semantic embeddings together. By the virtue of subsequent fusion block, the results of multi-level fusion can be furtherly enriched and fused. Experiments show that our model could achieve promising results on various datasets. Ablation study suggests that our method is well suited for zero shot learning.
Zhe Kong, Xin Wang 0086, Neng Gao, Yuhan Liu 0012, Chenyang Tu
ICMI1
2022 Fingerprint Presentation Attack Detection by Channel-Wise Feature Denoising
abstract
Due to the diversity of attack materials, fingerprint recognition systems (AFRSs) are vulnerable to malicious attacks. It is thus important to propose effective fingerprint presentation attack detection (PAD) methods for the safety and reliability of AFRSs. However, current PAD methods often exhibit poor robustness under new attack types settings. This paper thus proposes a novel channel-wise feature denoising fingerprint PAD (CFD-PAD) method by handling the redundant noise information ignored in previous studies. The proposed method learns important features of fingerprint images by weighing the importance of each channel and identifying discriminative channels and “noise” channels. Then, the propagation of “noise” channels is suppressed in the feature map to reduce interference. Specifically, a PA-Adaptation loss is designed to constrain the feature distribution to make the feature distribution of live fingerprints more aggregate and that of spoof fingerprints more disperse. Experimental results evaluated on the LivDet 2017 dataset showed that the proposed CFD-PAD can achieve 2.53% average classification error (ACE) and a 93.83% true detection rate when the false detection rate equals 1.0% (TDR@FDR=1%). Also, the proposed method markedly outperforms the best single-model-based methods in terms of ACE (2.53% vs. 4.56%) and TDR@FDR=1%(93.83% vs. 73.32%), which demonstrates its effectiveness. Although we have achieved a comparable result with the state-of-the-art multiple-model-based methods, there still is an increase in TDR@FDR=1% from 91.19% to 93.83%. In addition, the proposed model is simpler, lighter and more efficient and has achieved a 74.76% reduction in computation time compared with the state-of-the-art multiple-model-based method.The source code is available athttps://github.com/kongzhecn/cfd-pad.
Feng Liu 0013, Zhe Kong, Wentian Zhang, LinLin Shen
IEEE Trans. Inf. Forensics Secur.2