EDBT 2026 Demo / reviewers in the wild / expert
Qi Li 0005
dblp:181/2688-5
· DBLP profile ↗
46ranked-venue papers
9as first author
35since 2021 · last 2026
0000-0002-7905-2860ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 9 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 17 since 2021Security and privacy · 11 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks (Abstract Reprint)abstractDeep learning-based face recognition models are vulnerable to adversarial attacks. In contrast to general noises, the presence of imperceptible adversarial noises can lead to catastrophic errors in deep face recognition models. The primary difference between adversarial noise and general noise lies in its specificity. Adversarial attack methods give rise to noises tailored to the characteristics of the individual image and recognition model at hand. Diverse samples and recognition models can engender specific adversarial noise patterns, which pose significant challenges for adversarial defense. Addressing this challenge in the realm of face recognition presents a more formidable endeavor due to the inherent nature of face recognition as an open set task. In order to tackle this challenge, it is imperative to employ customized processing for each individual input sample. Drawing inspiration from the biological immune system, which can identify and respond to various threats, this paper aims to create an artificial immune system to provide adversarial defense for face recognition. The proposed defense model incorporates the principles of antibody cloning, mutation, selection, and memory mechanisms to generate a distinct antibody for each input sample, wherein the term antibody refers to a specialized noise removal manner. Furthermore, we introduce a self-supervised adversarial training mechanism that serves as a simulated rehearsal of immune system invasions. Extensive experimental results demonstrate the efficacy of the proposed method, surpassing state-of-the-art adversarial defense methods. Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
AAAI | 6 |
| 2026 | UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and PerceptionabstractThe remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual and textual modalities, especially in scenarios involving complex semantic instructions. However, existing approaches often rely heavily on vision-language models (VLMs) or modular designs for semantic guidance, leading to fragmented architectures and computational inefficiency. To address these challenges, we propose UniAlignment, a unified multimodal generation framework within a single diffusion transformer. UniAlignment introduces a dual-stream diffusion training strategy that incorporates both intrinsic-modal semantic alignment and cross-modal semantic alignment, thereby enhancing the model's cross-modal consistency and instruction-following robustness. Additionally, we present SemGen-Bench, a new benchmark specifically designed to evaluate multimodal semantic consistency under complex textual instructions. Extensive experiments across multiple tasks and benchmarks demonstrate that UniAlignment outperforms existing baselines, underscoring the significant potential of diffusion models in unified multimodal generation. Xinyang Song, Weining Wang 0001, Shaozhen Liu, Jingdong Chen, Qi Li 0005, Zhenan Sun |
AAAI | 7 |
| 2026 | Learning Unknown Spoof Prompts for Generalized Face Anti-Spoofing Using Only Real Face Images
Fangling Jiang, Qi Li 0005, Weining Wang 0001, Zhenan Sun |
Int. J. Comput. Vis. | 2 |
| 2026 | CAS-AIR-3D: A Large-scale Low-quality Multi-modal Face Database
Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Zhenan Sun, Tieniu Tan, Caifeng Shan |
Int. J. Comput. Vis. | 1 |
| 2026 | Learning Knowledge-Based Prompts for Robust 3D Mask Presentation Attack Detectionabstract3D mask presentation attack detection is crucial for protecting face recognition systems against the rising threat of 3D mask attacks. While most existing methods utilize multimodal features or remote photoplethysmography (rPPG) signals to distinguish between real faces and 3D masks, they face significant challenges, such as the high costs associated with multimodal sensors and limited generalization ability. Detection-related text descriptions offer concise, universal information and are cost-effective to obtain. However, the potential of vision-language multimodal features for 3D mask presentation attack detection remains unexplored. In this paper, we propose a novel knowledge-based prompt learning framework to explore the strong generalization capability of vision-language models for 3D mask presentation attack detection. Specifically, our approach incorporates entities and triples from knowledge graphs into the prompt learning process, generating fine-grained, task-specific explicit prompts that effectively harness the knowledge embedded in pre-trained vision-language models. Furthermore, considering different input images may emphasize distinct knowledge graph elements, we introduce a visual-specific knowledge filter based on an attention mechanism to refine relevant elements according to the visual context. Additionally, we leverage causal graph theory insights into the prompt learning process to further enhance the generalization ability of our method. During training, a spurious correlation elimination paradigm is employed, which removes category-irrelevant local image patches using guidance from knowledge-based text features, fostering the learning of generalized causal prompts that align with category-relevant local patches. Experimental results demonstrate that the proposed method achieves state-of-the-art intra- and cross-scenario detection performance on benchmark datasets. Fangling Jiang, Qi Li 0005, Weining Wang 0001, Caifeng Shan, Zhenan Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | HDTPose: A Hierarchical Decoding Transformer for End-to-End Single-Stage Multi-Person Pose EstimationabstractThe estimation performance of human keypoints varies significantly across different keypoint types. Compared to prominent joints such as the head and shoulders, smaller and more flexible limb joints present greater challenges in identification and localization. Current single-stage methods typically treat all body joints uniformly, overlooking the inherent differences and structural relationships among keypoints. To address this limitation, we propose HDTPose, a fully end-to-end multi-person pose estimation framework based on a hierarchical decoding transformer. HDTPose formulates multi-person pose estimation as a hierarchical set prediction problem and employs a hierarchical decoder to progressively decode joints in an end-to-end manner. The decoding process consists of two stages: an instance-aware stage and a joint-aware stage, which explicitly model instance-wise and joint-wise relationships, respectively. Additionally, we introduce a structure-guided joint attention mechanism that leverages kinematic relationships to refine pose predictions. Extensive experiments on the COCO and MPII benchmarks demonstrate that HDTPose outperforms existing state-of-the-art single-stage methods, underscoring the effectiveness and superiority of our approach. Wei Zhang 0255, Qi Li 0005, Zhenan Sun |
IJCNN | 3 |
| 2025 | FingerVeinSyn-5M: A Million-Scale Dataset and Benchmark for Finger Vein RecognitionabstractA major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we introduce FVeinSyn, a synthetic generator capable of producing diverse finger vein patterns with rich intra-class variations. Using FVeinSyn, we created FingerVeinSyn-5M -- the largest available finger vein dataset -- containing 5 million samples from 50,000 unique fingers, each with 100 variations including shift, rotation, scale, roll, varying exposure levels, skin scattering blur, optical blur, and motion blur. FingerVeinSyn-5M is also the first to offer fully annotated finger vein images, supporting deep learning applications in this field. Models pretrained on FingerVeinSyn-5M and fine-tuned with minimal real data achieve an average 53.91% performance gain across multiple benchmarks. The dataset is publicly available at: https://github.com/EvanWang98/FingerVeinSyn-5M. Yifan Wang 0036, Jie Gui, Baosheng Yu, Qi Li 0005, Zhenan Sun, Juho Kannala, Guoying Zhao 0001 |
ACM Multimedia | 4 |
| 2025 | AnyFace++: A Unified Framework for Free-Style Text-to-Face Synthesis and ManipulationabstractHuman faces contain rich semantic information that could hardly be described without a large vocabulary and complex sentence patterns. However, most existing text-to-image synthesis methods could only generate meaningful results based on limited sentence templates with words contained in the training set, which heavily impairs the generalization ability of these models. In this paper, we define a novel 'free-style' text-to-face generation and manipulation problem, and propose an effective solution, named AnyFace++, which is applicable to a much wider range of open-world scenarios. The CLIP model is involved in AnyFace++ for learning an aligned language-vision feature space, which also expands the range of acceptable vocabulary as it is trained on a large-scale dataset. To further improve the granularity of semantic alignment between text and images, a memory module is incorporated to convert the description with arbitrary length, format, and modality into regularized latent embeddings representing discriminative attributes of the target face. Moreover, the diversity and semantic consistency of generation results are improved by a novel semi-supervised training scheme and a series of newly proposed objective functions. Compared to state-of-the-art methods, AnyFace++ is capable of synthesizing and manipulating face images based on more flexible descriptions and producing realistic images with higher diversity. Jianxin Sun 0003, Qiyao Deng, Qi Li 0005, Muyi Sun, Yunfan Liu 0001, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Enhancing Adversarial Transferability With Alignment NetworkabstractDeep neural networks (DNNs) have been confirmed to exhibit vulnerability, as they are susceptible to deception by adversarial examples. Transfer-based attacks perturb a surrogate model and use the transferability of adversarial examples to attack other models. The effectiveness of these attacks relies heavily on the surrogate model, which often focuses on non-critical regions like backgrounds or object edges, leading to poor transferability. The intrinsic properties of the surrogate model fundamentally determine the performance of transfer-based attacks, yet this aspect has rarely been the focus of research. Therefore, we respectively design image masking operations for Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), forcing the model to reallocate attention to the critical regions. The attention of the surrogate model on the masked image and the original image is then aligned by inserting an alignment network inside the model. The modified surrogate model becomes more proficient in capturing the critical regions within the image, thereby generating more powerful adversarial examples. The proposed alignment network can be integrated into existing transfer-based attacks, significantly enhancing their performance. In addition, we also propose a novel feature-level attack based on the aligned attention, demonstrating superior performance compared to existing state-of-the-art feature-level attacks. Qi Li 0005, Yiwei Ru, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Biphasic Face Photo-Sketch Synthesis via Semantic-Driven Generative Adversarial Network With Graph Representation LearningabstractBiphasic face photo-sketch synthesis has significant practical value in wide-ranging fields such as digital entertainment and law enforcement. Previous approaches directly generate the photo-sketch in a global view, they always suffer from the low quality of sketches and complex photograph variations, leading to unnatural and low-fidelity results. In this article, we propose a novel semantic-driven generative adversarial network to address the above issues, cooperating with graph representation learning. Considering that human faces have distinct spatial structures, we first inject class-wise semantic layouts into the generator to provide style-based spatial information for synthesized face photographs and sketches. In addition, to enhance the authenticity of details in generated faces, we construct two types of representational graphs via semantic parsing maps upon input faces, dubbed the intraclass semantic graph (IASG) and the interclass structure graph (IRSG). Specifically, the IASG effectively models the intraclass semantic correlations of each facial semantic component, thus producing realistic facial details. To preserve the generated faces being more structure-coordinated, the IRSG models interclass structural relations among every facial component by graph representation learning. To further enhance the perceptual quality of synthesized images, we present a biphasic interactive cycle training strategy by fully taking advantage of the multilevel feature consistency between the photograph and sketch. Extensive experiments demonstrate that our method outperforms the state-of-the-art competitors on the CUHK Face Sketch (CUFS) and CUHK Face Sketch FERET (CUFSF) datasets. Xingqun Qi, Muyi Sun, Zijian Wang 0009, Jiaming Liu 0003, Qi Li 0005, Fang Zhao 0006, Shanghang Zhang, Caifeng Shan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | DCAPose: Improve One-Stage Multi-Person Pose Estimation with Dynamic Center AssignmentabstractSingle-stage methods for multi-person pose estimation have gained significant attention for their ability to concurrently localize person positions and perceive body structure in a single processing step. However, existing single-stage methods often rely on hand-crafted centers to represent the position of human instances. Such a simplification tends to overlook the intricate structure of the human body, resulting in a misalignment between the designated centers and the actual instance centers, which consequently leads to suboptimal performance. In this paper, we introduce DCAPose, a straightforward yet powerful pipeline to address this issue by redefining the process of center selection as a set prediction problem. Rather than directly supervising the center positions of instances, our approach considers each position on the center map as a potential instance candidate. We utilize a skeleton-aware bipartite matching loss to facilitate one-to-one matching between the poses of the candidate set and the ground truths. Additionally, we introduce a novel bidirectional hierarchical body representation to capture human structural information more accurately. Our method eliminates the need for Non-Maximum Suppression, greatly simplifying the processing pipeline and enabling end-to-end optimization. Extensive testing on challenging benchmarks COCO and CrowdPose confirms that DCAPose surpasses other leading single-stage methods, demonstrating the effectiveness and superiority of our framework. Wei Zhang 0255, Huiru Xie, Qi Li 0005, Zhenan Sun |
FG | 3 |
| 2024 | Open-Set Single-Domain Generalization for Robust Face Anti-Spoofing
Fangling Jiang, Qi Li 0005, Weining Wang 0001, Zhenan Sun |
Int. J. Comput. Vis. | 2 |
| 2024 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks
Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
Int. J. Comput. Vis. | 6 |
| 2024 | Learning Disentangled Representation for One-Shot Progressive Face SwappingabstractAlthough face swapping has attracted much attention in recent years, it remains a challenging problem. Existing methods leverage a large number of data samples to explore the intrinsic properties of face swapping without considering the semantic information of face images. Moreover, the representation of the identity information tends to be fixed, leading to suboptimal face swapping. In this paper, we present a simple yet efficient method named FaceSwapper, for one-shot face swapping based on Generative Adversarial Networks. Our method consists of a disentangled representation module and a semantic-guided fusion module. The disentangled representation module comprises an attribute encoder and an identity encoder, which aims to achieve the disentanglement of the identity and attribute information. The identity encoder is more flexible, and the attribute encoder contains more attribute details than its competitors. Benefiting from the disentangled representation, FaceSwapper can swap face images progressively. In addition, semantic information is introduced into the semantic-guided fusion module to control the swapped region and model the pose and expression more accurately. Experimental results show that our method achieves state-of-the-art results on benchmark datasets with fewer training samples. Qi Li 0005, Weining Wang 0001, Cheng-Zhong Xu 0001, Zhenan Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | r-FACE: Reference guided face component editing
Qiyao Deng, Jie Cao 0002, Yunfan Liu 0001, Qi Li 0005, Zhenan Sun |
Pattern Recognit. | 4 |
| 2024 | Produce Once, Utilize Twice for Anomaly DetectionabstractVisual anomaly detection aims at classifying and locating the regions that deviate from the normal appearance. Embedding-based methods and reconstruction-based methods are two main approaches for this task. The embedding-based methods typically predict the anomaly by measuring the distances between the deep representations of the test samples and a limited number of nominal samples, which enables these methods to be efficient but struggle in providing a fine-grained pixel-level anomaly location. The reconstruction-based methods rely on the pixel-level reconstruction errors to locate the anomaly, thereby the anomaly predictions are fine-grained. However, there are repetitive feature extractions and usually extra modules to guarantee the quality of the reconstructed images, resulting in unsatisfactory detection efficiency. In a nutshell, the prior methods are either not efficient or not precise enough for the industrial detection. To deal with this problem, we derive POUTA (Produce Once Utilize Twice for Anomaly detection), which improves both the accuracy and efficiency by reusing the discriminant information potential in the reconstructive network. We observe that the encoder and decoder representations of the reconstructive network are able to stand for the features of the original and reconstructed image respectively. And the discrepancies between the symmetric reconstructive representations provides roughly accurate anomaly information. To refine this information, a coarse-to-fine process is proposed in POUTA, which calibrates the semantics of each discriminative layer by the high-level representations and supervision loss. Equipped with the above modules, POUTA is endowed with the ability to provide a more precise anomaly location than the prior arts. Besides, the representation reusage also enables to exclude the feature extraction process in the discriminative network, which reduces the parameters and improves the efficiency. Extensive experiments show that, POUTA is superior or comparable to the prior methods with even less cost. Furthermore, POUTA also achieves better performance than the state-of-the-art few-shot anomaly detection methods without any special design, showing that POUTA has strong ability to learn representations inherent in the training data. Shuyuan Wang, Qi Li 0005, Huiyuan Luo, Chengkan Lv, Zhengtao Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Cross-Scenario Unknown-Aware Face Anti-Spoofing With Evidential Semantic Consistency LearningabstractIn recent years, domain adaptation techniques have been widely used to adapt face anti-spoofing models to a cross-scenario target domain. Most previous methods assume that the Presentation Attack Instruments (PAIs) in such cross-scenario target domain are same as in the source domain. However, as the malicious users are free to use any form of unknown PAIs to attack the system, this assumption does not always hold in practical applications of face anti-spoofing. Thus, unknown PAIs would inevitably lead to significant performance degradation, since samples of known and unknown PAIs usually have large differences. In this paper, we propose an Evidential Semantic Consistency Learning (ESCL) framework to address this problem. Specifically, a regularized evidential deep learning strategy with a two-way balance of class probability and uncertainty is leveraged to produce uncertainty scores for unknown PAI detection. Meanwhile, entropy optimization-based semantic consistency learning strategy is also employed to encourage features of live and known PAIs to be gathered in the label-conditioned clusters across the source and target domains, while make the features of unknown PAIs to be self-clustered according to intrinsic semantic information. In addition, a new evaluation metric, KUHAR, is proposed to comprehensively evaluate the error rate of known classes and unknown PAIs. Extensive experimental results on six public datasets demonstrate the effectiveness of our method in generalizing face anti-spoofing models to both known classes and unknown PAIs with different types and quantities in a cross-scenario testing domain. Our method achieves state-of-the-art performance on eight different protocols. Fangling Jiang, Yunfan Liu 0001, Haolin Si, Jingjing Meng, Qi Li 0005 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Improving Transferability of Adversarial Samples via Critical Region-Oriented Feature-Level AttackabstractDeep neural networks (DNNs) have received a lot of attention because of their impressive progress in computer vision. However, it has been recently shown that DNNs are vulnerable to being spoofed by carefully crafted adversarial samples. These samples are generated by specific attack algorithms that can obfuscate the target model without being detected by humans. Recently, feature-level attacks have been the focus of research due to their high transferability. Existing state-of-the-art feature-level attacks all improve the transferability by greedily changing the attention of the model. However, for images that contain multiple target class objects, the attention of different models may differ significantly. Thus greedily changing attention may cause the adversarial samples corresponding to these images to fall into the local optimum of the surrogate model. Furthermore, due to the great structural differences between vision transformers (ViTs) and convolutional neural networks (CNNs), adversarial samples generated on CNNs with feature-level attacks are more difficult to successfully attack ViTs. To overcome these drawbacks, we perform the Critical Region-oriented Feature-level Attack (CRFA) in this paper. Specifically, we first propose the Perturbation Attention-aware Weighting (PAW), which destroys critical regions of the image by performing feature-level attention weighting on the adversarial perturbations without changing the model attention as much as possible. Then we propose the Region ViT-critical Retrieval (RVR), which enables the generator to accommodate the transferability of adversarial samples on ViTs by adding extra prior knowledge of ViTs to the decoder. Extensive experiments demonstrate significant performance improvements achieved by our approach, i.e., improving the fooling rate by 19.9% against CNNs and 25.0% against ViTs as compared to state-of-the-art feature-level attack method. Qi Li 0005, Fangling Jiang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological IntelligenceabstractAffective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP. Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun |
ACM Multimedia | 6 |
| 2023 | Adversarial Learning Domain-Invariant Conditional Features for Robust Face Anti-spoofing
Fangling Jiang, Qi Li 0005, Zhenan Sun |
Int. J. Comput. Vis. | 2 |
| 2023 | GAN-Based Facial Attribute ManipulationabstractFacial Attribute Manipulation (FAM) aims to aesthetically modify a given face image to render desired attributes, which has received significant attention due to its broad practical applications ranging from digital entertainment to biometric forensics. In the last decade, with the remarkable success of Generative Adversarial Networks (GANs) in synthesizing realistic images, numerous GAN-based models have been proposed to solve FAM with various problem formulation approaches and guiding information representations. This paper presents a comprehensive survey of GAN-based FAM methods with a focus on summarizing their principal motivations and technical details. The main contents of this survey include: (i) an introduction to the research background and basic concepts related to FAM, (ii) a systematic review of GAN-based FAM methods in three main categories, and (iii) an in-depth discussion of important properties of FAM methods, open issues, and future research directions. This survey not only builds a good starting point for researchers new to this field but also serves as a reference for the vision community. Yunfan Liu 0001, Qi Li 0005, Qiyao Deng, Zhenan Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Semantic-based conditional generative adversarial hashing with pairwise labels
Qi Li 0005, Weining Wang 0001, Yuan Yan Tang, Cheng-Zhong Xu 0001, Zhenan Sun |
Pattern Recognit. | 1 |
| 2023 | Towards Spatially Disentangled Manipulation of Face Images With Pre-Trained StyleGANsabstractGenerative Adversarial Networks with style-based generators could successfully synthesize realistic images from input latent code. Moreover, recent studies have revealed that interpretable translations of generated images could be obtained by linearly traversing in the latent space. However, in most existing latent spaces, linear interpolation often leads to ‘spatially entangled modification’ in the manipulation result, which is undesirable in many real-world applications where local editing is required. To solve this problem, we propose to manipulate the latent code in the ‘style space’ and analyze its advantage in achieving spatial disentanglement. Furthermore, we point out the weakness of simply interpolating in the style space and propose ‘Style Intervention’, a lightweight optimization-based algorithm, to further improve the visual fidelity of manipulation results. The performance of our method is verified with the task of attribute editing on high-resolution face images. Both qualitative and quantitative results demonstrate the advantage of image translation in the style space and the effectiveness of our method on both real and synthetic images. Yunfan Liu 0001, Qi Li 0005, Qiyao Deng, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Semantic-Aware Noise Driven Portrait Synthesis and ManipulationabstractSemantic portrait synthesis has drawn consistent attention and has made significant progress, yet achieving style diversity and semantic controllability simultaneously is still a challenge. Existing methods either 1) directly take a semantic label map as input, ignoring various possibilities of semantic styles, or 2) sample global noise as input, ignoring controllability of local semantics. To fill this gap, we propose semantic-aware noise, a simple but effective input that tackles both issues and shows improved results over baselines. Semantic-aware noise introduces semantic information into noise, and each semantic is sampled from the noise separately, combining the semantic controllability and the noise sampling diversity. To further expand and manipulate real images, we propose a novel ternary network structure, allowing simultaneous diverse semantic image synthesis and real image manipulation in a unified framework. Extensive experiments demonstrate that the proposed method achieves quantitatively superior and perceptually pleasing results compared to state-of-the-art methods. We also analyze the performance of our method with respect to different noise structures and real-life applications in diverse synthesis, interactive manipulation, and extreme pose scenarios. Qiyao Deng, Qi Li 0005, Jie Cao 0002, Yunfan Liu 0001, Zhenan Sun |
IEEE Trans. Multim. | 2 |
| 2022 | AnyFace: Free-style Text-to-Face Synthesis and ManipulationabstractExisting text-to-image synthesis methods generally are only applicable to words in the training dataset. However, human faces are so variable to be described with limited words. So this paper proposes the first free-style text-to-face method namely AnyFace enabling much wider open world applications such as metaverse, social media, cosmetics, forensics, etc. AnyFace has a novel two-stream framework for face image synthesis and manipulation given arbitrary descriptions of the human face. Specifically, one stream performs text-to-face generation and the other conducts face image reconstruction. Facial text and image features are extracted using the CLIP (Contrastive Language-Image Pre-training) encoders. And a collaborative Cross Modal Distillation (CMD) module is designed to align the linguistic and visual features across these two streams. Furthermore, a Diverse Triplet Loss (DT loss) is developed to model fine-grained features and improve facial diversity. Extensive experiments on Multi-modal CelebA-HQ and CelebAText-HQ demonstrate significant advantages of AnyFace over state-of-the-art methods. AnyFace can achieve high-quality, high-resolution, and high-diversity face synthesis and manipulation results without any constraints on the number and content of input captions. Jianxin Sun 0003, Qiyao Deng, Qi Li 0005, Muyi Sun, Zhenan Sun |
CVPR | 3 |
| 2022 | MOST-Net: A Memory Oriented Style Transfer Network for Face Sketch SynthesisabstractFace sketch synthesis has been widely used in multimedia entertainment and law enforcement. Despite the recent developments in deep neural networks, accurate and realistic face sketch synthesis is still a challenging task due to the diversity and complexity of human faces. Current image-to-image translation-based face sketch synthesis frequently encounters over-fitting problems when it comes to small-scale datasets. To tackle this problem, we present an end-to-end Memory Oriented Style Transfer Network (MOST-Net) for face sketch synthesis which can produce high-fidelity sketches with limited data. Specifically, an external self-supervised dynamic memory module is introduced to capture the domain alignment knowledge in the long term. In this way, our proposed model could obtain the domain-transfer ability by establishing the durable relationship between faces and corresponding sketches on the feature level. Furthermore, we design a novel Memory Refinement Loss (MR Loss) for feature alignment in the memory module, which enhances the accuracy of memory slots in an unsupervised manner. Extensive experiments on the CUFS and the CUFSF datasets show that our MOST-Net achieves state-of-the-art performance, especially in terms of the Structural Similarity Index(SSIM). Fan Ji, Muyi Sun, Xingqun Qi, Qi Li 0005, Zhenan Sun |
ICPR | 4 |
| 2021 | One Shot Face Swapping on MegapixelsabstractFace swapping has both positive applications such as entertainment, human-computer interaction, etc., and negative applications such as DeepFake threats to politics, economics, etc. Nevertheless, it is necessary to understand the scheme of advanced methods for high-quality face swapping and generate enough and representative face swapping images to train DeepFake detection algorithms. This paper proposes the first Megapixel level method for one shot Face Swapping (or MegaFS for short). Firstly, MegaFS organizes face representation hierarchically by the proposed Hierarchical Representation Face Encoder (HieRFE) in an extended latent space to maintain more facial details, rather than compressed representation in previous face swapping methods. Secondly, a carefully designed Face Transfer Module (FTM) is proposed to transfer the identity from a source image to the target by a non-linear trajectory without explicit feature disentanglement. Finally, the swapped faces can be synthesized by StyleGAN2 with the benefits of its training stability and powerful generative capability. Each part of MegaFS can be trained separately so the requirement of our model for GPU memory can be satisfied for megapixel face swapping. In summary, complete face representation, stable training, and limited memory usage are the three novel contributions to the success of our method. Extensive experiments demonstrate the superiority of MegaFS and the first megapixel level face swapping database is released for research on DeepFake detection and face image editing in the public domain. Yuhao Zhu 0003, Qi Li 0005, Cheng-Zhong Xu 0001, Zhenan Sun |
CVPR | 2 |
| 2021 | CAS-AIR-3D Face: A Low-Quality, Multi-Modal and Multi-Pose 3D Face DatabaseabstractBenefiting from deep learning with large scale face databases, 2D face recognition has made significant progress in recent years. However, it still highly depends on lighting conditions and human poses, and suffers from face spoofing problem. In contrast, 3D face recognition reveals a new path that can overcome the previous limitations of 2D face recognition. One of the most important problems for 3D face recognition is to construct a suitable database, which can be exploited to train different 3D face recognition algorithms. In this work, we propose a new database, CAS-AIR-3D Face, for low-quality 3D face recognition. It includes 24713 videos from 3093 individuals, which is captured by Intel RealSense SR305. The database contains three modalities: color, depth and near infrared, and is rich in pose, expression, occlusion and distance variations. To the best of our konwledge, CAS-AIR-3D Face is the largest low-quality 3D face database in terms of the number of individuals and the sample variations. Moreover, we preprocess the data via a sophisticated face alignment method, and Point Cloud Spherical Cropping Method (SCM) is leveraged to remove the background noise in the depth images. Finally, an evaluation protocol is designed for fair comparison, and extensive experiments are conducted with different backbone networks to provide different baselines on this database. Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Caifeng Shan |
IJCB | 1 |
| 2021 | DFGC 2021: A DeepFake Game CompetitionabstractThis paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang |
IJCB | 7 |
| 2021 | Face Sketch Synthesis via Semantic-Driven Generative Adversarial NetworkabstractFace sketch synthesis has made significant progress with the development of deep neural networks in these years. The delicate depiction of sketch portraits facilitates a wide range of applications like digital entertainment and law enforcement. However, accurate and realistic face sketch generation is still a challenging task due to the illumination variations and complex backgrounds in the real scenes. To tackle these challenges, we propose a novel Semantic-Driven Generative Adversarial Network (SDGAN) which embeds global structure-level style injection and local class-level knowledge re-weighting. Specifically, we conduct facial saliency detection on the input face photos to provide overall facial texture structure, which could be used as a global type of prior information. In addition, we exploit face parsing layouts as the semantic-level spatial prior to enforce globally structural style injection in the generator of SDGAN. Furthermore, to enhance the realistic effect of the details, we propose a novel Adaptive Re-weighting Loss (ARLoss) which dedicates to balance the contributions of different semantic classes. Experimentally, our extensive experiments on CUFS and CUFSF datasets show that our proposed algorithm achieves state-of-the-art performance. Xingqun Qi, Muyi Sun, Weining Wang 0001, Xiaoxiao Dong, Qi Li 0005, Caifeng Shan |
IJCB | 5 |
| 2021 | Bita-Net: Bi-temporal Attention Network for Facial Video Forgery DetectionabstractDeep forgery detection on video data has attracted remarkable research attention in recent years due to its potential in defending forgery attacks. However, existing methods either only focus on the visual evidence within individual images, or are too sensitive to fluctuations across frames. To address these issues, this paper propose a novel model, named Bita-Net, to detect forgery faces in video data. The network design of Bita-Net is inspired by the mechanism of how human beings detect forgery data, i.e. browsing and scrutinizing, which is reflected by the two-pathway architecture of Bita-Net. Concretely, the browsing pathway scans the entire video at a high frame rate to check the temporal consistency, while the scrutinizing pathway focuses on analyzing key frames of the video at a lower frame rate. Furthermore, an attention branch is introduced to improve the forgery detection ability of the scrutinizing pathway. Extensive experiment results demonstrate the effectiveness and generalization ability of Bita-Net on various popular face forensics detection datasets, including FaceForensics++, CelebDF, DeepfakeTIMIT and UADFV. Yiwei Ru, Yunfan Liu 0001, Jianxin Sun 0003, Qi Li 0005 |
IJCB | 5 |
| 2021 | Multi-caption Text-to-Face Synthesis: Dataset and AlgorithmabstractText-to-Face synthesis with multiple captions is still an important yet less addressed problem because of the lack of effective algorithms and large-scale datasets. We accordingly propose a Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images. With a novel Sentence Features Injection Module, SEA-T2F can integrate any number of captions into the network. In addition, an attention mechanism named Attention for Multiple Captions is proposed to fuse multiple word features and synthesize fine-grained details. Considering text-to-face generation is an ill-posed problem, we also introduce an attribute loss to guide the network to generate sentence-related attributes. Existing datasets for text-to-face are either too small or roughly generated according to attribute labels, which is not enough to train deep learning based methods to synthesize natural face images. Therefore, we build a large-scale dataset named CelebAText-HQ, in which each image is manually annotated with 10 captions. Extensive experiments demonstrate the effectiveness of our algorithm. Jianxin Sun 0003, Qi Li 0005, Weining Wang 0001, Jian Zhao 0006, Zhenan Sun |
ACM Multimedia | 2 |
| 2021 | Boosting End-to-end Multi-Object Tracking and Person Search via Knowledge DistillationabstractMulti-Object Tracking (MOT) and Person Search both demand to localize and identify specific targets from raw image frames. Existing methods can be classified into two categories, namely two-step strategy and end-to-end strategy. Two-step approaches have high accuracy but suffer from costly computations, while end-to-end methods show greater efficiency with limited performance. In this paper, we dissect the gap between two-step and end-to-end strategy and propose a simple yet effective end-to-end framework with knowledge distillation. Our proposed framework is simple in concept and easy to benefit from external datasets. Experimental results demonstrate that our model performs competitively with other sophisticated two-step and end-to-end methods in multi-object tracking and person search. Wei Zhang 0255, Lingxiao He, Xingyu Liao, Wu Liu 0005, Qi Li 0005, Zhenan Sun |
ACM Multimedia | 6 |
| 2021 | Controllable Multi-Attribute Editing of High-Resolution Face ImagesabstractIn recent years, significant progress has been achieved in face image editing due to the success of Generative Adversarial Network (GAN). However, state-of-the-art face editing methods mainly suffer from the following two limitations: 1) they are only applicable to face images with relative low-resolutions and 2) multi-attribute face editing may generate uncontrollable changes in non-target face attribute categories. To solve these problems, we propose a novel High-Quality Generative Adversarial Network (HQ-GAN) for controllable editing of multiple face attributes in high-resolution images. HQ-GAN has two novel ideas to break the limitations of resolution and controllability correspondingly: 1) fine-grained textures and realistic details of high-resolution face images are better preserved with the aid of textural features extracted by the wavelet transform module and 2) desired multi-attribute targets of face editing are emphasized using a weighted binary cross-entropy (BCE) loss so that the influence on non-target attributes is greatly reduced. To the best of our knowledge, HQ-GAN is the first attempt to achieve continuous editing of multiple face attributes on high-resolution images of the CelebA-HQ using only 28 000 training samples. Extensive qualitative results demonstrate the superiority of the proposed method in rendering realistic high-resolution face images with accurate attribute modification, and comprehensive quantitative results show that the proposed method significantly outperforms state-of-the-art face editing methods. Qiyao Deng, Qi Li 0005, Jie Cao 0002, Yunfan Liu 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | A3GAN: An Attribute-Aware Attentive Generative Adversarial Network for Face AgingabstractFace aging has received significant research attention in recent years. Although great progress has been achieved with the success of Generative Adversarial Networks (GANs) in synthesizing realistic images, most existing GAN-based face aging methods have two main problems: 1) unnatural changes of high-level semantic information due to the insufficient consideration of prior knowledge of input faces, and 2) distortions of low-level image content (e.g. modifications in age-irrelevant regions). In this article, we introduce A3GAN, an Attribute-Aware Attentive face aging model to address the above issues. Facial attribute vectors are regarded as the conditional information and embedded into both the generator and discriminator, encouraging synthesized faces to be faithful to attributes of corresponding inputs. To improve the visual fidelity of generation results, we leverage the attention mechanism to restrict modifications to age-related areas and preserve image details. Unlike previous works with attention modules, we introduce face parsing maps to help the generator distinguish image regions of interest and suppress attention activation elsewhere. Moreover, the wavelet packet transform is employed to capture textural features at multiple scales in the frequency space. Extensive experimental results demonstrate the effectiveness of our model in synthesizing photo-realistic aged face images and achieving state-of-the-art performance on popular datasets. Yunfan Liu 0001, Qi Li 0005, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Age Progression and Regression with Spatial Attention ModulesabstractAge progression and regression refers to aesthetically rendering a given face image to present effects of face aging and rejuvenation, respectively. Although numerous studies have been conducted in this topic, there are two major problems: 1) multiple models are usually trained to simulate different age mappings, and 2) the photo-realism of generated face images is heavily influenced by the variation of training images in terms of pose, illumination, and background. To address these issues, in this paper, we propose a framework based on conditional Generative Adversarial Networks (cGANs) to achieve age progression and regression simultaneously. Particularly, since face aging and rejuvenation are largely different in terms of image translation patterns, we model these two processes using two separate generators, each dedicated to one age changing process. In addition, we exploit spatial attention mechanisms to limit image modifications to regions closely related to age changes, so that images with high visual fidelity could be synthesized for in-the-wild cases. Experiments on multiple datasets demonstrate the ability of our model in synthesizing lifelike face images at desired ages with personalized features well preserved, and keeping age-irrelevant regions unchanged. Qi Li 0005, Yunfan Liu 0001, Zhenan Sun |
AAAI | 1 |
| 2020 | Reference Guided Face Component EditingabstractFace portrait editing has achieved great progress in recent years. However, previous methods either 1) operate on pre-defined face attributes, lacking the flexibility of controlling shapes of high-level semantic facial components (e.g., eyes, nose, mouth), or 2) take manually edited mask or sketch as an intermediate representation for observable changes, but such additional input usually requires extra efforts to obtain. To break the limitations (e.g. shape, mask or sketch) of the existing methods, we propose a novel framework termed r FACE (Reference Guided FAce Component Editing) for diverse and controllable face component editing with geometric changes. Specifically, r-FACE takes an image inpainting model as the backbone, utilizing reference images as conditions for controlling the shape of face components. In order to encourage the framework to concentrate on the target face components, an example-guided attention module is designed to fuse attention features and the target face component features extracted from the reference image. Through extensive experimental validation and comparisons, we verify the effectiveness of the proposed framework. Qiyao Deng, Jie Cao 0002, Yunfan Liu 0001, Zhenhua Chai, Qi Li 0005, Zhenan Sun |
IJCAI | 5 |
| 2020 | A General Framework for Deep Supervised Discrete Hashing
Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 1 |
| 2019 | Attribute-Aware Face Aging With Wavelet-Based Generative Adversarial NetworksabstractSince it is difficult to collect face images of the same subject over a long range of age span, most existing face aging methods resort to unpaired datasets to learn age mappings. However, the matching ambiguity between young and aged face images inherent to unpaired training data may lead to unnatural changes of facial attributes during the aging process, which could not be solved by only enforcing identity consistency like most existing studies do. In this paper, we propose an attribute-aware face aging model with wavelet based Generative Adversarial Networks (GANs) to address the above issues. To be specific, we embed facial attribute vectors into both the generator and discriminator of the model to encourage each synthesized elderly face image to be faithful to the attribute of its corresponding input. In addition, a wavelet packet transform (WPT) module is incorporated to improve the visual fidelity of generated images by capturing age-related texture details at multiple scales in the frequency space. Qualitative results demonstrate the ability of our model in synthesizing visually plausible face images, and extensive quantitative evaluation results show that the proposed method achieves state-of-the-art performance on existing datasets. Yunfan Liu 0001, Qi Li 0005, Zhenan Sun |
CVPR | 2 |
| 2019 | Adversarial Learning Semantic Volume for 2D/3D Face Shape Regression in the WildabstractRegression based methods have revolutionized 2D landmark localization with the exploitation of deep neural networks and massive annotated datasets in the wild. However, it remains challenging for 3D landmark localization due to the lack of annotated datasets and the ambiguous nature of landmarks under 3D perspective. This paper revisits regression based methods and proposes an adversarial voxel and coordinate regression framework for 2D and 3D facial landmark localization in real-world scenarios. First, a semantic volumetric representation is introduced to encode the per-voxel likelihood of positions being the 3D landmarks. Then, an end-to-end pipeline is designed to jointly regress the proposed volumetric representation and the coordinate vector. Such a pipeline not only enhances the robustness and accuracy of the predictions but also unifies the 2D and 3D landmark localization so that 2D and 3D datasets could be utilized simultaneously. Further, an adversarial learning strategy is exploited to distill 3D structure learned from synthetic datasets to real-world datasets under weakly supervised settings, where an auxiliary regression discriminator is proposed to encourage the network to produce plausible predictions for both synthetic and real-world images. The effectiveness of our method is validated on benchmark datasets 3DFAW and AFLW2000-3D for both 2D and 3D facial landmark localization tasks. Experimental results show that the proposed method achieves significant improvements over previous state-of-the-art methods. Hongwen Zhang 0001, Qi Li 0005, Zhenan Sun |
IEEE Trans. Image Process. | 2 |
| 2018 | Global and Local Consistent Age Generative Adversarial NetworksabstractAge progression/regression is a challenging task due to the complicated and non-linear transformation in human aging process. Many researches have shown that both global and local facial features are essential for face representation [1], but previous GAN based methods mainly focused on the global feature in age synthesis. To utilize both global and local facial information, we propose a Global and Local Consistent Age Generative Adversarial Network (GLCA-GAN). In our generator, a global network learns the whole facial structure and simulates the aging trend of the whole face, while three crucial facial patches are progressed or regressed by three local networks aiming at imitating subtle changes of crucial facial subregions. To preserve most of the details in age-attribute-irrelevant areas, our generator learns the residual face. Moreover, we employ an identity preserving loss to better preserve the identity information, as well as age preserving loss to enhance the accuracy of age synthesis. A pixel loss is also adopted to preserve detailed facial information of the input face. Our proposed method is evaluated on three face aging datasets, i.e., CACD dataset, Morph dataset and FG-NET dataset. Experimental results show appealing performance of the proposed method by comparing with the state-of-the-art. Peipei Li 0002, Yibo Hu 0001, Qi Li 0005, Ran He 0001, Zhenan Sun |
ICPR | 3 |
| 2018 | Joint Voxel and Coordinate Regression for Accurate 3D Facial Landmark Localizationabstract3D face shape is more expressive and viewpoint-consistent than its 2D counterpart. However, 3D facial landmark localization in a single image is challenging due to the ambiguous nature of landmarks under 3D perspective. Existing approaches typically adopt a suboptimal two-step strategy, performing 2D landmark localization followed by depth estimation. In this paper, we propose the Joint Voxel and Coordinate Regression (JVCR) method for 3D facial landmark localization, addressing it more effectively in an end-to-end fashion. First, a compact volumetric representation is proposed to encode the per-voxel likelihood of positions being the 3D landmarks. The dimensionality of such a representation is fixed regardless of the number of target landmarks, so that the curse of dimensionality could be avoided. Then, a stacked hourglass network is adopted to estimate the volumetric representation from coarse to fine, followed by a 3D convolution network that takes the estimated volume as input and regresses 3D coordinates of the face shape. In this way, the 3D structural constraints between landmarks could be learned by the neural network in a more efficient manner. Moreover, the proposed pipeline enables end-to-end training and improves the robustness and accuracy of 3D facial landmark localization. The effectiveness of our approach is validated on the 3DFAW and AFLW2000-3D datasets. Experimental results show that the proposed method achieves state-of-the-art performance in comparison with existing methods. Hongwen Zhang 0001, Qi Li 0005, Zhenan Sun |
ICPR | 2 |
| 2018 | Combining Data-Driven and Model-Driven Methods for Robust Facial Landmark DetectionabstractFacial landmark detection is an important yet challenging task for real-world computer vision applications. This paper proposes an effective and robust approach for facial landmark detection by combining data- and model-driven methods. First, a fully convolutional network (FCN) is trained to compute response maps of all facial landmark points. Such a data-driven method could make full use of holistic information in a facial image for global estimation of facial landmarks. After that, the maximum points in the response maps are fitted with a pre-trained point distribution model (PDM) to generate the initial facial shape. This model-driven method is able to correct the inaccurate locations of outliers by considering the shape prior information. Finally, a weighted version of regularized landmark mean-shift (RLMS) is employed to fine-tune the facial shape iteratively. This estimation-correction-tuning process perfectly combines the advantages of the global robustness of the data-driven method (FCN), outlier correction capability of the model-driven method (PDM), and non-parametric optimization of RLMS. Results of extensive experiments demonstrate that our approach achieves state-of-the-art performances on challenging data sets, including 300W, AFLW, AFW, and COFW. The proposed method is able to produce satisfying detection results on face images with exaggerated expressions, large head poses, and partial occlusions. Hongwen Zhang 0001, Qi Li 0005, Zhenan Sun, Yunfan Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Fast multi-view face alignment via multi-task auto-encodersabstractFace alignment is an important problem in computer vision. It is still an open problem due to the variations of facial attributes (e.g., head pose, facial expression, illumination variation). Many studies have shown that face alignment and facial attribute analysis are often correlated. This paper develops a two-stage multi-task Auto-encoders framework for fast face alignment by incorporating head pose information to handle large view variations. In the first and second stages, multi-task Auto-encoders are used to roughly locate and further refine facial landmark locations with related pose information, respectively. Besides, the shape constraint is naturally encoded into our two-stage face alignment framework to preserve facial structures. A coarse-to-fine strategy is adopted to refine the facial landmark results with the shape constraint. Furthermore, the computational cost of our method is much lower than its deep learning competitors. Experimental results on various challenging datasets show the effectiveness of the proposed method. Qi Li 0005, Zhenan Sun, Ran He 0001 |
IJCB | 1 |
| 2017 | Deep Supervised Discrete HashingabstractWith the rapid growth of image and video data on the web, hashing has been extensively studied for image or video search in recent years. Benefiting from recent advances in deep learning, deep hashing methods have achieved promising results for image retrieval. However, there are some limitations of previous deep hashing methods (e.g., the semantic information is not fully exploited). In this paper, we develop a deep supervised discrete hashing algorithm based on the assumption that the learned binary codes should be ideal for classification. Both the pairwise label information and the classification information are used to learn the hash codes within one stream framework. We constrain the outputs of the last layer to be binary codes directly, which is rarely investigated in deep hashing algorithm. Because of the discrete nature of hash codes, an alternating minimization method is used to optimize the objective function. Experimental results have shown that our method outperforms current state-of-the-art methods on benchmark datasets. Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan |
NIPS | 1 |
| 2016 | Transformation invariant subspace clustering
Qi Li 0005, Zhenan Sun, Zhouchen Lin, Ran He 0001, Tieniu Tan |
Pattern Recognit. | 1 |