Ran He 0001

dblp:61/6198-1 · DBLP profile ↗
← Back
242ranked-venue papers
25as first author
116since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 181 · 22 first-author · 87 since 2021Graphics, computer vision, multimedia, augmented reality and games · 130 · 10 first-author · 56 since 2021Security and privacy · 23 · 16 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 5Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 CoGrad3D: Spatially-Coupled Timestep Optimization with Orthogonal Gradient Fusion for 3D Generation
abstract
Score Distillation Sampling has driven recent advances in text-to-3D generation. However, current approaches often fail to produce 3D assets that are both rich in detail and consistent across viewpoints. These limitations primarily arise from imbalanced guidance on fine-grained details and an overdependence on single-view optimization—issues exacerbated by the excessive randomness in selecting diffusion timesteps and camera configurations. Such deficiencies commonly lead to blurry textures and inter-view inconsistencies, which degrade visual realism and hinder practical deployment. To tackle these challenges, we introduce CoGrad3D, a unified generative refinement framework that adopts a continuously adaptive optimization strategy. By dynamically modulating the optimization focus based on real-time convergence signals, CoGrad3D ensures balanced progress toward both geometric completeness and high-fidelity detail. Concretely, we propose an adaptive region sampling strategy that emphasizes under-converged viewing areas, promoting stable and uniform optimization. To facilitate the transition from coarse geometry to fine-grained reconstruction, we develop a region-aware temporal scheduling scheme that integrates global training dynamics with local convergence feedback. Furthermore, we introduce a gradient fusion mechanism that consolidates historical gradients from adjacent viewpoints, mitigating view-specific artifacts and promoting the emergence of coherent 3D structures. Extensive experiments demonstrate that CoGrad3D substantially surpasses existing methods in both geometric consistency and texture fidelity, enabling the generation of high-quality, view-consistent 3D models from textual descriptions.
Haoyang Tong, Jin Liu 0040, Jie Cao 0002, Ran He 0001
AAAI6
2026 What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
abstract
Test-Time Reinforcement Learning (TTRL) enables Large Language Models (LLMs) to enhance reasoning capabilities on unlabeled test streams by deriving pseudo-rewards from majority voting consensus.However, existing TTRL methods rely exclusively on positive pseudo-labeling strategies.Such reliance becomes vulnerable under challenging scenarios where answer distributions are highly dispersed, resulting in weak consensus that inadvertently reinforces incorrect trajectories as supervision signals.In this paper, we propose SCRL (Selective-Complementary Reinforcement Learning), a robust test-time reinforcement learning framework that effectively mitigates label noise amplification.SCRL develops Selective Positive Pseudo-Labeling, which enforces strict consensus criteria to filter unreliable majorities.Complementarily, SCRL introduces Entropy-Gated Negative Pseudo-Labeling, the first negative supervision mechanism in TTRL, to reliably prune incorrect trajectories based on generation uncertainty.Extensive experiments on multiple reasoning benchmarks demonstrate that SCRL achieves substantial improvements over baselines, while maintaining robust generalization and training stability under constrained rollout budgets.Our code is available at https://github.com/Jasper
Jian Liang 0001, Yanbo Wang 0004, Shuo Lu, Ran He 0001, Tieniu Tan
ACL (1)5
2026 Uncertainty-Aware Source-Free Adaptive Image Restoration with State Space Augmentation
Yuang Ai, Jie Cao 0002, Ran He 0001, Huaibo Huang
Int. J. Comput. Vis.3
2026 Advancing Vision Transformer With Enhanced Spatial Priors
abstract
In recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and suffers from quadratic computational complexity, limiting its applicability. To address these issues, we have proposed RMT, a robust vision backbone with explicit spatial priors for general purposes. RMT utilizes Manhattan distance decay to introduce spatial information and employs a horizontal and vertical decomposition attention method to model global information. Building on the strengths of RMT, Euclidean enhanced Vision Transformer (EVT) is an expanded version that incorporates several key improvements. Firstly, EVT uses a more reasonable Euclidean distance decay to enhance the modeling of spatial information, allowing for a more accurate representation of spatial relationships compared to the Manhattan distance used in RMT. Secondly, EVT abandons the decomposed attention mechanism featured in RMT and instead adopts a simpler spatially-independent grouping approach, providing the model with greater flexibility in controlling the number of tokens within each group. By addressing these modifications, EVT offers a more sophisticated and adaptable approach to incorporating spatial priors into the Self-Attention mechanism, thus overcoming some of the limitations associated with RMT and further enhancing its applicability in various computer vision tasks. Extensive experiments on Image Classification, Object Detection, Instance Segmentation, and Semantic Segmentation demonstrate that EVT exhibits exceptional performance. Without additional training data, EVT achieves 86.6% top1-acc on ImageNet-1 k.
Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Hongmin Liu 0001, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Bidirectional intervention attention network for audio-visual matching
Jiaxiang Wang 0001, Aihua Zheng, Dequan Li, Chenglong Li 0002, Wenjuan Cheng, Ran He 0001
Pattern Recognit.6
2026 InfoBFR: Real-World Blind Face Restoration via Information Bottleneck
abstract
Previous blind face restoration (BFR) methods have primarily leveraged facial priors from pretrained GAN or diffusion models. These neural BFR models suffer from diverse neural degradations, such as prior bias, topological distortion, textural distortion, and artifact residues, which limit their real-world generalization in complex real-world scenarios. In this paper, we propose an effective framework,InfoBFR, to address neural degradation from an information-theoretic perspective, which achieves BFR boosting in diverse wild and heterogeneous scenes. Specifically, on the basis of the results from pretrained BFR models, InfoBFR considers information compression by using a manifold information bottleneck (MIB) and manifold information compensation (MIC) with efficient diffusion LoRA to conduct information optimization. InfoBFR effectively synthesizes high-fidelity faces with texture and structure boosting. Comprehensive experimental results demonstrate the high boosting performance of InfoBFR (nearly 82%) for state-of-the-art GAN-based and diffusion-based BFR methods, as it can complete BFR tasks in approximately 70 ms and has 4M trainable parameters. It is promising that InfoBFR is the first unified postprocessing restorer universally employed by diverse BFR models to overcome the limitations of neural degradation.
Nan Gao 0001, Jia Li 0044, Huaibo Huang, Ran He 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Exploring Vacant Classes in Label-Skewed Federated Learning
abstract
Label skews, characterized by disparities in local label distribution across clients, pose a significant challenge in federated learning. As minority classes suffer from worse accuracy due to overfitting on local imbalanced data, prior methods often incorporate class-balanced learning techniques during local training. Although these methods improve the mean accuracy across all classes, we observe that vacant classes—referring to categories absent from a client's data distribution—remain poorly recognized. Besides, there is still a gap in the accuracy of local models on minority classes compared to the global model. This paper introduces FedVLS, a novel approach to label-skewed federated learning that integrates both vacant-class distillation and logit suppression simultaneously. Specifically, vacant-class distillation leverages knowledge distillation during local training on each client to retain essential information related to vacant classes from the global model. Moreover, logit suppression directly penalizes network logits for non-label classes, effectively addressing misclassifications in minority classes that may be biased toward majority classes. Extensive experiments validate the efficacy of FedVLS, demonstrating superior performance compared to previous state-of-the-art (SOTA) methods across diverse datasets with varying degrees of label skews.
Kuangpu Guo, Yuhe Ding, Jian Liang 0001, Zilei Wang, Ran He 0001, Tieniu Tan
AAAI5
2025 Protecting Model Adaptation from Trojans in the Unlabeled Data
abstract
Model adaptation tackles the distribution shift problem with a pre-trained model instead of raw data, which has become a popular paradigm due to its great privacy protection. Existing methods always assume adapting to a clean target domain, overlooking the security risks of unlabeled samples. This paper for the first time explores the potential trojan attacks on model adaptation launched by well-designed poisoning target data. Concretely, we provide two trigger patterns with two poisoning strategies for different prior knowledge owned by attackers. These attacks achieve a high success rate while maintaining the normal performance on clean samples in the test stage. To defend against such backdoor injection, we propose a plug-and-play method named DiffAdapt, which can be seamlessly integrated with existing adaptation algorithms. Experiments across commonly used benchmarks and adaptation methods demonstrate the effectiveness of DiffAdapt. We hope this work will shed light on the safety of transfer learning with unlabeled data.
Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
AAAI3
2025 Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
abstract
Recently, scaling test-time compute on Large Language Models (LLM) has garnered wide attention. However, there has been limited investigation of how various reasoning prompting strategies perform as scaling. In this paper, we focus on a standard and realistic scaling setting: majority voting. We systematically conduct experiments on 6 LLMs $\times$ 8 prompting strategies $\times$ 6 benchmarks. Experiment results consistently show that as the sampling time and computational overhead increase, complicated prompting strategies with superior initial performance gradually fall behind simple Chain-of-Thought. We analyze this phenomenon and provide theoretical proofs. Additionally, we propose a probabilistic method to efficiently predict scaling performance and identify the best prompting strategy under large sampling times, eliminating the need for resource-intensive inference processes in practical applications. Furthermore, we introduce two ways derived from our theoretical analysis to significantly improve the scaling performance. We hope that our research can promote to re-examine the role of complicated prompting, unleash the potential of simple prompting strategies, and provide new insights for enhancing test-time scaling performance. Code is available at https://github.com/MraDonkey/rethinking_prompting.
Yexiang Liu, Zekun Li 0007, Zhi Fang, Nan Xu 0014, Ran He 0001, Tieniu Tan
ACL (1)5
2025 Breaking the Low-Rank Dilemma of Linear Attention
abstract
The Softmax attention mechanism in Transformer models is notoriously computationally expensive due to its quadratic complexity, posing significant challenges in vision applications. In contrast, linear attention offers a far more efficient solution by reducing the complexity to linear levels. However, linear attention often suffers significant performance degradation compared to Softmax attention. Our experiments indicate that this performance drop stems from the low-rank nature of linear attention’s output feature map, which hinders its ability to adequately model complex spatial information. To address this low-rank dilemma, we conduct rank analysis from two perspectives: the KV buffer and the output features. Consequently, we introduce Rank-Augmented Linear Attention (RALA), which rivals the performance of Softmax attention while maintaining linear complexity and high efficiency. Building upon RALA, we construct the Rank-Augmented Vision Linear Transformer (RAVLT). Extensive experiments demonstrate that RAVLT achieves excellent performance across various vision tasks. Specifically, without using any additional labels, data, or supervision during training, RAVLT achieves an 84.4% Top-1 accuracy on ImageNet-1k with only 26M parameters and 4.6G FLOPs. This result significantly surpasses previous linear attention mechanisms, fully illustrating the potential of RALA. Code will be available at https://github.com/qhfan/RALA.
Qihang Fan, Huaibo Huang, Ran He 0001
CVPR3
2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
abstract
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual data is still insufficiently explored, highlighting the lack of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, and reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models with an average accuracy of 75%, compared to 71.9% for GPT-4o. The results also demonstrate that Video-MME is a universal benchmark that applies to both image and video MLLMs. Further analysis indicates that subtitle and audio information could significantly enhance video understanding. Besides, a decline in MLLM performance is observed as video duration increases for all models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data, shedding light on future MLLM development. Project page: https://video-mme.github.io.
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Shuhuai Ren, Renrui Zhang, Yunhang Shen, Mengdan Zhang, Peixian Chen, Shaohui Lin, Sirui Zhao, Ke Li 0015, Tong Xu 0001, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He 0001, Xing Sun 0001
CVPR20
2025 R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
abstract
Vision-language models (VLMs), such as CLIP, have gained significant popularity as foundation models, with numerous fine-tuning methods developed to enhance performance on downstream tasks. However, due to their inherent vulnerability and the common practice of selecting from a limited set of open-source models, VLMs suffer from a higher risk of adversarial attacks than traditional vision models. Existing defense techniques typically rely on adversarial fine-tuning during training, which requires labeled data and lacks of flexibility for downstream tasks. To address these limitations, we propose robust test-time prompt tuning (R-TPT), which mitigates the impact of adversarial attacks during the inference stage. We first reformulate the classic marginal entropy objective by eliminating the term that introduces conflicts under adversarial conditions, retaining only the pointwise entropy minimization. Furthermore, we introduce a plug-and-play reliability-based weighted ensembling strategy, which aggregates useful information from reliable augmented views to strengthen the defense. R-TPT enhances defense against adversarial attacks without requiring labeled training data while offering high flexibility for inference tasks. Extensive experiments on widely used benchmarks with various attacks demonstrate the effectiveness of R-TPT. The code is available in https://github.com/TomSheng21/R-TPT.
Lijun Sheng, Jian Liang 0001, Zilei Wang, Ran He 0001
CVPR4
2025 Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
abstract
Multi-modal large language models (MLLMs) have made significant progress, yet their safety alignment remains limited. Typically, current open-source MLLMs rely on the alignment inherited from their language module to avoid harmful generations. However, the lack of safety measures specifically designed for multi-modal inputs creates an alignment gap, leaving MLLMs vulnerable to vision-domain attacks such as typographic manipulation. Current methods utilize a carefully designed safety dataset to enhance model defense capability, while the specific knowledge or patterns acquired from the high-quality dataset remain unclear. Through comparison experiments, we find that the alignment gap primarily arises from data distribution biases, while image content, response quality, or the contrastive behavior of the dataset makes little contribution to boosting multi-modal safety. To further investigate this and identify the key factors in improving MLLM safety, we propose finetuning MLLMs on a small set of benign instruct-following data with responses replaced by simple, clear rejection sentences. Experiments show that, without the need for labor-intensive collection of high-quality malicious data, model safety can still be significantly improved, as long as a specific fraction of rejection data exists in the finetuning set, indicating the security alignment is not lost but rather obscured during multi-modal pretraining or instruction finetuning. Simply correcting the underlying data bias could narrow the safety gap in the vision domain. Warning: This paper contains harmful images and AI-generated contents which may be offensive.
Yanbo Wang 0004, Jiyang Guan, Jian Liang 0001, Ran He 0001
CVPR4
2025 DM-DPR: Diffusion and Mamba-based Degradation Prediction for Blind Face Restoration
abstract
Blind Face Restoration (BFR), which involves converting low-quality facial images with unknown and varied degradation into high-quality counterparts, suffers from issues of sub-optimal restoration and over-correction due to inconsistent degradation levels. To rectify the above issue, inspired by the State Space model, especially the improved version Mamba’s enhanced long-range dependencies modeling ability and Stable Diffusion’s ability in integrating multi-modal prompts, we introduce a novel approach, Diff-Mamba Degradation Prediction Restoration (DM-DPR), to leverage the combination of a Mamba prompt generation framework with Stable Diffusion-based image restoration. Its core lies in two primary components: a Mamba-based multi-modal prompt generator that quantifies the degradation severity, generating corresponding textual and visual prompts, additionally with a multi-modal prompt driven Stable Diffusion process that adjusts restoration efforts based on the estimated degradation level. Derived from the CelebA-Test, we create degraded datasets exhibiting a wide range of degradation severity. Extensive experimental evaluations demonstrate that DM-DPR substantially surpasses existing state-of-the-art methods, thereby robustly establishing its enhanced capability to manage varying degrees of image degradation.
Guorong Yuan, Huaibo Huang, Jie Cao 0002, Yuang Ai, Ran He 0001
FG6
2025 Growing to Detect: A Dynamic Prototype Tree with Structured Replay for Incremental Deepfake Detection
abstract
The rapid advancement of deepfake technology poses significant threats to social trust. Recent research has improved detectors by adapting to emerging deepfakes using a limited number of samples through incremental learning. However, these approaches often overlook the scarcity of novel samples, resulting in insufficient learning of forgery patterns. To overcome this challenge, we propose a Replay-based Dynamic Prototype Network that integrates two key modules: the Dynamic Prototype Tree (DPT) module and the Similarity Subtree Replay (SSR) strategy. The DPT module dynamically introduces prototypes through a hierarchical tree structure to effectively adapt to new deepfakes. It expands prototypes based on similarity, thereby retaining the knowledge learned from previous prototypes while learning new forgery patterns. The SSR strategy mitigates catastrophic forgetting by stabilizing learned features through the replay of relevant subtrees. Experimental results demonstrate that our approach outperforms existing methods across five datasets, particularly on high-quality face swap samples generated by diffusion-based methods, achieving an AUC of 85.86% on the cross-dataset task from FaceForensics++ to DiffSwap.
Junxian Duan, Jie Cao 0002, Aihua Zheng, Ran He 0001
IJCB5
2025 Rectifying Magnitude Neglect in Linear Attention
abstract
As the core operator of Transformers, Softmax Attention exhibits excellent global modeling capabilities. However, its quadratic complexity limits its applicability to vision tasks. In contrast, Linear Attention shares a similar formulation with Softmax Attention while achieving linear complexity, enabling efficient global information modeling. Nevertheless, Linear Attention suffers from a significant performance degradation compared to standard Softmax Attention. In this paper, we analyze the underlying causes of this issue based on the formulation of Linear Attention. We find that, unlike Softmax Attention, Linear Attention entirely disregards the magnitude information of the Query. This prevents the attention score distribution from dynamically adapting as the Query scales. As a result, despite its structural similarity to Softmax Attention, Linear Attention exhibits a significantly different attention score distribution. Based on this observation, we propose Magnitude-Aware Linear Attention (MALA), which modifies the computation of Linear Attention to fully incorporate the Query's magnitude. This adjustment allows MALA to generate an attention score distribution that closely resembles Softmax Attention while exhibiting a more well-balanced structure. We evaluate the effectiveness of MALA on multiple tasks, including image classification, object detection, instance segmentation, semantic segmentation, natural language processing, speech recognition, and image generation. Our MALA achieves strong results on all of these tasks. Code will be available at https://github.com/qhfan/MALA
Qihang Fan, Huaibo Huang, Yuang Ai, Ran He 0001
ICCV4
2025 Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Ran He 0001
ICCV4
2025 Cooperative Pseudo Labeling for Unsupervised Federated Classification
Kuangpu Guo, Lijun Sheng, Yongcan Yu, Jian Liang 0001, Zilei Wang, Ran He 0001
ICCV6
2025 Towards Robust Defense Against Customization via Protective Perturbation Resistant to Diffusion-based Purification
abstract
Diffusion models like Stable Diffusion have become prominent in visual synthesis tasks due to their powerful customization capabilities, which also introduce significant security risks, including deepfakes and copyright infringement. In response, a class of methods known as protective perturbation emerged, which mitigates image misuse by injecting imperceptible adversarial noise. However, purification can remove protective perturbations, thereby exposing images again to the risk of malicious forgery. In this work, we formalize the anti-purification task, highlighting challenges that hinder existing approaches, and propose a simple diagnostic protective perturbation named AntiPure. AntiPure exposes vulnerabilities of purification within the "purification-customization" workflow, owing to two guidance mechanisms: 1) Patch-wise Frequency Guidance, which reduces the model's influence over high-frequency components in the purified image, and 2) Erroneous Timestep Guidance, which disrupts the model's denoising strategy across different timesteps. With additional guidance, AntiPure embeds imperceptible perturbations that persist under representative purification settings, achieving effective post-customization distortion. Experiments show that, as a stress test for purification, AntiPure achieves minimal perceptual discrepancy and maximal distortion, outperforming other protective perturbation methods within the purification-customization workflow.
Wenkui Yang, Jie Cao 0002, Junxian Duan, Ran He 0001
ICCV4
2025 Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent Debate
abstract
Large Language Models (LLMs) have seen significant progress but continue to struggle with persistent reasoning mistakes. Previous methods of *self-reflection* have been proven limited due to the models’ inherent fixed thinking patterns. While Multi-Agent Debate (MAD) attempts to mitigate this by incorporating multiple agents, it often employs the same reasoning methods, even though assigning different personas to models. This leads to a "fixed mental set", where models rely on homogeneous thought processes without exploring alternative perspectives. In this paper, we introduce Diverse Multi-Agent Debate (DMAD), a method that encourages agents to think with distinct reasoning approaches. By leveraging diverse problem-solving strategies, each agent can gain insights from different perspectives, refining its responses through discussion and collectively arriving at the optimal solution. DMAD effectively breaks the limitations of fixed mental sets. We evaluate DMAD against various prompting techniques, including *self-reflection* and traditional MAD, across multiple benchmarks using both LLMs and Multimodal LLMs. Our experiments show that DMAD consistently outperforms other methods, delivering better results than MAD in fewer rounds. Code is available at https://github.com/MraDonkey/DMAD.
Yexiang Liu, Jie Cao 0002, Zekun Li 0001, Ran He 0001, Tieniu Tan
ICLR4
2025 LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
abstract
Low-rank adaptation, also known as LoRA, has emerged as a prominent method for parameter-efficient fine-tuning of foundation models. Despite its computational efficiency, LoRA still yields inferior performance compared to full fine-tuning. In this paper, we first uncover a fundamental connection between the optimization processes of LoRA and full fine-tuning: using LoRA for optimization is mathematically equivalent to full fine-tuning using a low-rank gradient for parameter updates. And this low-rank gradient can be expressed in terms of the gradients of the two low-rank matrices in LoRA. Leveraging this insight, we introduce LoRA-Pro, a method that enhances LoRA's performance by strategically adjusting the gradients of these low-rank matrices. This adjustment allows the low-rank gradient to more accurately approximate the full fine-tuning gradient, thereby narrowing the performance gap between LoRA and full fine-tuning. Furthermore, we theoretically derive the optimal solutions for adjusting the gradients of the low-rank matrices, applying them during fine-tuning in LoRA-Pro. We conduct extensive experiments across natural language understanding, dialogue generation, mathematical reasoning, code generation, and image classification tasks, demonstrating that LoRA-Pro substantially improves LoRA's performance, effectively narrowing the gap with full fine-tuning. Our code is publicly available at https://github.com/mrflogs/LoRA-Pro.
Zhengbo Wang, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
ICLR3
2025 Degradation-Aware Multi-Task Image Restoration with State Space Models
abstract
Image restoration (IR) has made significant strides, evolving from basic pixel-wise restoration to more advanced techniques capable of handling diverse degradations. However, current all-in-one IR methods face challenges in effectively generalizing across complex real-world degradation scenarios. To address this challenge, we propose multi-task Restoration Mamba (ReMamba), a novel framework that leverages the power of state space models for long-sequence modeling to extract fine-grained degradation features from low-quality data. Guided by degradation-type prediction, which helps the model accurately identify and differentiate between various degradation types in the input, ReMamba employs prompt learning across both textual and visual modalities and enhances the restoration capabilities of downstream stable diffusion networks by injecting prompts as prior knowledge. Extensive experiments on both synthetic and real-world datasets, covering six distinct IR tasks, demonstrate the adaptability, generalizability, and robustness of ReMamba, highlighting its effectiveness in addressing a wide range of degradation types under challenging conditions.
Purui Bai, Huaibo Huang, Jie Cao 0002, Yuang Ai, Ran He 0001
ICME6
2025 MTSD: Simple Yet Effective Self-Distillation for Generalizable Deepfake Detection
abstract
The rapid advancement of Deepfake technology necessitates detection systems with strong generalization capabilities. Existing methods often depend on architectural modifications or dataset-specific prior knowledge, which limits their scalability and practical deployment in real-world scenarios. We propose Multi-Teacher Self-Distillation (MTSD), a simple yet effective and generalizable strategy to enhance model generalization. MTSD comprises two key steps. First, diverse teacher generation leverages independently trained teacher models with varying dataset sampling sequences to capture complementary decision boundaries. Second, self-distillation feature fusion integrates these diverse features using a cross-attention mechanism, allowing the student model to approximate an ideal feature distribution for improved generalization. This strategy avoids architectural changes and dataset-specific adjustments, ensuring simplicity in implementation and deployment. Moreover, the multi-teacher generation and feature fusion steps are discarded after training, preserving computational efficiency during inference. Experimental results demonstrate that MTSD significantly improves model generalization, offering a practical and scalable solution for Deepfake detection.
Dexu Zhu, Jie Cao 0002, Jiangnan Shao, Junxian Duan, Ran He 0001
ICME6
2025 DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
abstract
Diffusion Transformer (DiT), a promising diffusion model for visual generation, demonstrates impressive performance but incurs significant computational overhead. Intriguingly, analysis of pre-trained DiT models reveals that global self-attention is often redundant, predominantly capturing local patterns—highlighting the potential for more efficient alternatives. In this paper, we revisit convolution as an alternative building block for constructing efficient and expressive diffusion models. However, naively replacing self-attention with convolution typically results in degraded performance. Our investigations attribute this performance gap to the higher channel redundancy in ConvNets compared to Transformers. To resolve this, we introduce a compact channel attention mechanism that promotes the activation of more diverse channels, thereby enhancing feature diversity. This leads to Diffusion ConvNet (DiCo), a family of diffusion models built entirely from standard ConvNet modules, offering strong generative performance with significant efficiency gains. On class-conditional ImageNet generation benchmarks, DiCo-XL achieves an FID of 2.05 at 256$\times$256 resolution and 2.53 at 512$\times$512, with a **2.7$\times$** and **3.1$\times$** speedup over DiT-XL/2, respectively. Furthermore, experimental results on MS-COCO demonstrate that the purely convolutional DiCo exhibits strong potential for text-to-image generation.
Yuang Ai, Qihang Fan, Xuefeng Hu, Zhenheng Yang, Ran He 0001, Huaibo Huang
NeurIPS5
2025 MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
abstract
Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization. The data are released at the project page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation.
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Jinrui Yang, Xiawu Zheng, Ke Li 0015, Xing Sun 0001, Yunsheng Wu, Rongrong Ji, Caifeng Shan, Ran He 0001
NeurIPS14
2025 VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
abstract
Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in both vision and speech tasks remains a challenge due to the fundamental modality differences. In this paper, we propose a carefully designed multi-stage training methodology that progressively trains LLM to understand both visual and speech information, ultimately enabling fluent vision and speech interaction. Our approach not only preserves strong vision-language capacity, but also enables efficient speech-to-speech dialogue capabilities without separate ASR and TTS modules, significantly accelerating multimodal end-to-end response speed. By comparing against state-of-the-art counterparts across benchmarks for image, video, and speech, we demonstrate that our omni model is equipped with both strong visual and speech capabilities, making omni understanding and interaction.
Chaoyou Fu, Haojia Lin, Yifan Zhang 0004, Yunhang Shen, Haoyu Cao 0001, Zuwei Long, Heting Gao, Ke Li 0015, Xiawu Zheng, Rongrong Ji, Xing Sun 0001, Caifeng Shan, Ran He 0001
NeurIPS16
2025 Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
abstract
The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce distinct safety risks. To bridge this gap, we introduce Video-SafetyBench, the first comprehensive benchmark designed to evaluate the safety of LVLMs under video-text attacks. It comprises 2,264 video-text pairs spanning 48 fine-grained unsafe categories, each pairing a synthesized video with either a harmful query, which contains explicit malice, or a benign query, which appears harmless but triggers harmful behavior when interpreted alongside the video. To generate semantically accurate videos for safety evaluation, we design a controllable pipeline that decomposes video semantics into subject images (what is shown) and motion text (how it moves), which jointly guide the synthesis of query-relevant videos. To effectively evaluate uncertain or borderline harmful outputs, we propose RJScore, a novel LLM-based metric that incorporates the confidence of judge models and human-aligned decision threshold calibration. Extensive experiments show that benign-query video composition achieves average attack success rates of 67.2%, revealing consistent vulnerabilities to video-induced attacks. We believe Video-SafetyBench will catalyze future research into video-based safety evaluation and defense strategies.
Xuannan Liu, Zekun Li 0001, Zheqi He, Peipei Li 0002, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang 0023, Ran He 0001
NeurIPS9
2025 The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
abstract
Test-time adaptation (TTA) methods have gained significant attention for enhancing the performance of vision-language models (VLMs) such as CLIP during inference, without requiring additional labeled data. However, current TTA researches generally suffer from major limitations such as duplication of baseline results, limited evaluation metrics, inconsistent experimental settings, and insufficient analysis. These problems hinder fair comparisons between TTA methods and make it difficult to assess their practical strengths and weaknesses. To address these challenges, we introduce TTA-VLM, a comprehensive benchmark for evaluating TTA methods on VLMs. Our benchmark implements 8 episodic TTA and 7 online TTA methods within a unified and reproducible framework, and evaluates them across 15 widely used datasets. Unlike prior studies focused solely on CLIP, we extend the evaluation to SigLIP—a model trained with a Sigmoid loss—and include training-time tuning methods such as CoOp, MaPLe, and TeCoA to assess generality. Beyond classification accuracy, TTA-VLM incorporates various evaluation metrics, including robustness, calibration, out-of-distribution detection, and stability, enabling a more holistic assessment of TTA methods. Through extensive experiments, we find that 1) existing TTA methods produce limited gains compared to the previous pioneering work; 2) current TTA methods exhibit poor collaboration with training-time fine-tuning methods; 3) accuracy gains frequently come at the cost of reduced model trustworthiness. We release TTA-VLM to provide fair comparison and comprehensive evaluation of TTA methods for VLMs, and we hope it encourages the community to develop more reliable and generalizable TTA strategies. The code is available in https://github.com/TomSheng21/tta-vlm.
Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
NeurIPS3
2025 ZeroPatcher: Training-free Sampler for Video Inpainting and Editing
abstract
Video inpainting and editing have long been challenging tasks in the video generation community, requiring extensive computational resources and large datasets to train models with satisfactory performance. Recent breakthroughs in large-scale video foundation models have greatly enhanced text-to-video generation capabilities. This naturally leads to the idea of leveraging the prior knowledge from these powerful generators to facilitate video inpainting and editing. In this work, we investigate the feasibility of employing pre-trained text-to-video foundation models for high-quality video inpainting and editing without additional training. Specifically, we introduce a model-agnostic denoising sampler that optimizes the trajectory by maximizing the log-likelihood expectation conditioned on the known video segments. To enable efficient dynamic object removal and replacement, we propose a latent mask fuser that performs accurate video masking directly in latent space, eliminating the need for explicit VAE decoding and encoding. We implement our approach in widely-used foundation generators such as CogVideoX and HunyuanVideo, demonstrating the model-agnostic nature of our sampler. Comprehensive quantitative and qualitative evaluations confirm that our method achieves outstanding video inpainting and editing performance in a plug-and-play fashion.
Shaoshu Yang, Yingya Zhang, Ran He 0001
NeurIPS3
2025 Frustratingly Easy Feature Reconstruction for Out-of-Distribution Detection
Yingsheng Wang, Shuo Lu, Jian Liang 0001, Aihua Zheng, Ran He 0001
PRCV (9)5
2025 Trustworthy forgery detection with causal inference
Junxian Duan, Fan Ji, Yi Li 0018, Ran He 0001
Sci. China Inf. Sci.5
2025 Test-time Forgery Detection with Spatial-Frequency Prompt Learning
Junxian Duan, Yuang Ai, Shenyuan Huang, Huaibo Huang, Jie Cao 0002, Ran He 0001
Int. J. Comput. Vis.7
2025 Sample Correlation for Fingerprinting Deep Face Recognition
Jiyang Guan, Jian Liang 0001, Yanbo Wang 0004, Ran He 0001
Int. J. Comput. Vis.4
2025 A Comprehensive Survey on Test-Time Adaptation Under Distribution Shifts
Jian Liang 0001, Ran He 0001, Tieniu Tan
Int. J. Comput. Vis.2
2025 Adaptive Interaction and Correction Attention Network for Audio-Visual Matching
abstract
Audio-visual matching techniques aim to recognize and match information across different identities by learning a similarity metric across modalities. However, modal differences arise from insufficient cross-modal correlations and noise interference, which substantially hinder the performance of traditional deep metric learning methods in audio-visual matching tasks. To address the modal differences issue, we propose a novel Adaptive Interactive and Correction Attention Network (AICANet). This network efficiently captures deep information connections, generating modality-consistent feature embeddings within a unified metric framework. The core of AICANet is its two-pronged approach to reducing modal differences. First, we propose the Adaptive Interactive Attention (AIA) module, which flexibly establishes associations among cross-modal local features using dynamically generated pseudo-labels. Second, we propose the Adaptive Correction Attention (ACA) mechanism, which employs an adaptive threshold to de-interference effectively and accurately adjust the representation of local feature associations. Notably, the ACA mechanism is suitable for both intra-modal and inter-modal refined attention correction. Additionally, we design a relative distance stretching metric loss (LRDSM), which reinforces the similarity invariance of feature embeddings in a uniform space and enhances matching accuracy. Extensive tests on the VoxCeleb and VoxCeleb2 datasets demonstrate that AICANet outperforms leading existing algorithms across several evaluation metrics, validating its superior performance. The codes can be found at https://github.com/w1018979952/AICANet.
Jiaxiang Wang 0001, Aihua Zheng, Lei Liu 0049, Chenglong Li 0002, Ran He 0001, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Uncertainty-Aware Bilateral Transformer for Accurate and Reliable Iris Segmentation
abstract
Iris segmentation is a deterministic and critical part of the iris recognition system. However, its performance is usually degraded by data uncertainty in acquisition and annotation, impeding more accurate recognition of the iris recognition system. In the paper, we propose a bilateral self-attention by exploring spatial and visual relationships to effectively distinguish between iris and non-iris regions, then design a bilateral Transformer by enhancing spatial perception and hierarchical feature fusion to mitigate the impact of acquisition uncertainty. Besides, iris segmentation uncertainty learning is developed to estimate the uncertainty map according to prediction discrepancy. With the estimated uncertainty, a weighting scheme and a regularization term are designed to minimize the effect of annotation uncertainty. To investigate data uncertainty, the paper presents a challenging near-infrared iris dataset named UTIris. It comprises 3,690 images with high acquisition uncertainty and provides rich segmentation masks to explore annotation uncertainty. Furthermore, we manually label a large-scale iris dataset, ND-0405 [1], with additional binary maps of iris masks to evaluate segmentation performance. Experimental results on UTIris and four other databases demonstrate the effectiveness of the proposed method in iris segmentation, and its segmentation improvement consequently promotes recognition accuracy.
Jianze Wei, Xingyu Gao 0001, Yunlong Wang 0003, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2024 Heterogeneous Test-Time Training for Multi-Modal Person Re-identification
abstract
Multi-modal person re-identification (ReID) seeks to mitigate challenging lighting conditions by incorporating diverse modalities. Most existing multi-modal ReID methods concentrate on leveraging complementary multi-modal information via fusion or interaction. However, the relationships among heterogeneous modalities and the domain traits of unlabeled test data are rarely explored. In this paper, we propose a Heterogeneous Test-time Training (HTT) framework for multi-modal person ReID. We first propose a Cross-identity Inter-modal Margin (CIM) loss to amplify the differentiation among distinct identity samples. Moreover, we design a Multi-modal Test-time Training (MTT) strategy to enhance the generalization of the model by leveraging the relationships in the heterogeneous modalities and the information existing in the test data. Specifically, in the training stage, we utilize the CIM loss to further enlarge the distance between anchor and negative by forcing the inter-modal distance to maintain the margin, resulting in an enhancement of the discriminative capacity of the ultimate descriptor. Subsequently, since the test data contains characteristics of the target domain, we adapt the MTT strategy to optimize the network before the inference by using self-supervised tasks designed based on relationships among modalities. Experimental results on benchmark multi-modal ReID datasets RGBNT201, Market1501-MM, RGBN300, and RGBNT100 validate the effectiveness of the proposed method. The codes can be found at https://github.com/ziwang1121/HTT.
Zi Wang 0013, Huaibo Huang, Aihua Zheng, Ran He 0001
AAAI4
2024 DeVAn: Dense Video Annotation for Video-Language Models
abstract
Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fang, Ding Zhou, Huaibo Huang, Ran He, Hongxia Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan, Huaibo Huang, Ran He 0001, Hongxia Yang
ACL (1)7
2024 Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration
abstract
Despite substantial progress, all-in-one image restoration (IR) grapples with persistent challenges in handling intricate real-world degradations. This paper introduces MPerceiver: a novel multimodal prompt learning approach that harnesses Stable Diffusion (SD) priors to enhance adaptiveness, generalizability and fidelity for all-in-one im-age restoration. Specifically, we develop a dual-branch module to master two types of SD prompts: textual for holistic representation and visual for multiscale detail rep-resentation. Both prompts are dynamically adjusted by degradation predictions from the CLIP image encoder, en-abling adaptive responses to diverse unknown degradations. Moreover, a plug-in detail refinement module im-proves restoration fidelity via direct encoder-to-decoder in-formation transformation. To assess our method, MPer-ceiver is trained on 9 tasks for all-in-one IR and outper-forms state-of-the-art task-specific methods across many tasks. Post multitask pre-training, MPerceiver attains a generalized representation in low-level vision, exhibiting remarkable zero-shot and few-shot capabilities in unseen tasks. Extensive experiments on 16 IR tasks underscore the superiority of MPerceiver in terms of adaptiveness, gener-alizability and fidelity.
Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, Ran He 0001
CVPR5
2024 Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation Transformer
abstract
Unsupervised Domain Adaptation (UDA) can effectively address domain gap issues in real-world image Super-Resolution (SR) by accessing both the source and target data. Considering privacy policies or transmission restrictions of source data in practical scenarios, we propose a SOurce-free Domain Adaptation framework for image SR (SODA-SR) to address this issue, i.e., adapt a source-trained model to a target domain with only unlabeled target data. SODA-SR leverages the source-trained model to generate refined pseudo-labels for teacher-student learning. To better utilize pseudo-labels, we propose a novel wavelet-based augmentation method, named Wavelet Augmentation Transformer (WAT), which can be flexibly incorporated with existing networks, to implicitly produce useful augmented data. WAT learns low-frequency information of varying levels across diverse samples, which is aggregated efficiently via deformable attention. Furthermore, an uncertainty-aware self-training mechanism is proposed to improve the accuracy of pseudo-labels, with inaccurate predictions being rectified by uncertainty estimation. To acquire better SR results and avoid overfitting pseudo-labels, several regularization losses are proposed to constrain target LR and SR images in the frequency domain. Experiments show that without accessing source data, SODA-SR outperforms state-of-the-art UDA methods in both synthetic→real and real→real adaptation settings, and is not constrained by specific network architectures.
Yuang Ai, Xiaoqiang Zhou, Huaibo Huang, Ran He 0001
CVPR5
2024 RMT: Retentive Networks Meet Vision Transformers
abstract
Vision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. How-ever, the core component of ViT, Self-Attention, lacks ex-plicit spatial priors and bears a quadratic computational complexity, thereby constraining the applicability of ViT. To alleviate these issues, we draw inspiration from the re-cent Retentive Network (RetNet) in the field of NLP, and propose RMT, a strong vision backbone with explicit spa-tial prior for general purposes. Specifically, we extend the RetNet's temporal decay mechanism to the spatial do-main, and propose a spatial decay matrix based on the Manhattan distance to introduce the explicit spatial prior to Self-Attention. Additionally, an attention decomposition form that adeptly adapts to explicit spatial prior is proposed, aiming to reduce the computational burden of modeling global information without disrupting the spa-tial decay matrix. Based on the spatial decay matrix and the attention decomposition form, we can flexibly integrate explicit spatial prior into the vision backbone with lin-ear complexity. Extensive experiments demonstrate that RMT exhibits exceptional performance across various vision tasks. Specifically, without extra training data, RMT achieves 84.8% and 86.1% top-l acc on ImageNet-lk with 27MI4.5GFLOPs and 96M/18.2GFLOPs. For downstream tasks, RMT achieves 54.5 box AP and 47.2 mask AP on the COCO detection task, and 52.8 mloU on the ADE20K se-mantic segmentation task.
Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Hongmin Liu 0001, Ran He 0001
CVPR5
2024 Backdoor Defense via Test-Time Detecting and Repairing
abstract
Deep neural networks have played a crucial part in many critical domains, such as autonomous driving, face recognition, and medical diagnosis. However, deep neural networks are facing security threats from backdoor attacks and can be manipulated into attacker-decided behaviors by the backdoor attacker. To defend the backdoor, prior research has focused on using clean data to remove backdoor attacks before model deployment. In this paper, we investigate the possibility of defending against backdoor attacks by utilizing test-time partially poisoned data to remove the backdoor from the model. To address the problem, a two-stage method TTBD is proposed. In the first stage, we propose a backdoor sample detection method DDP to identify poisoned samples from a batch of mixed, partially poisoned samples. Once the poisoned samples are detected, we employ Shapley estimation to calculate the contribution of each neuron's significance in the network, locate the poisoned neurons, and prune them to remove backdoor in the models. Our experiments demonstrate that TTBD removes the backdoor successfully with only a batch of partially poisoned data across different model architectures and datasets against different types of backdoor attacks.
Jiyang Guan, Jian Liang 0001, Ran He 0001
CVPR3
2024 STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
Yongcan Yu, Lijun Sheng, Ran He 0001, Jian Liang 0001
ECCV (81)3
2024 PortraitDAE: Line-Drawing Portraits Style Transfer from Photos via Diffusion Autoencoder with Meaningful Encoded Noise
abstract
The line-drawing portrait is a kind of highly abstract art that contains a sparse set of continuous graphical elements such as lines to capture a person's facial features. Due to their abstract artistic form, common style transfer methods fail to synthesize high-quality line-drawing portraits from photos. Previous works mostly concentrate on GANs, often requiring pre-calculated landmarks acquired by other models and using extra classifiers with complicated structures to capture local facial features. We propose a novel idea without these extra operations based on diffusion models, which is more flexible and stable than GAN-based methods. We utilize the diffusion-based decoder in the Diffusion Autoencoder to encode the input image to an encoded noise that contains much meaningful stochastic information by running the deterministic generative process backward. By fully utilizing the encoded noise, our method can effectively preserve the identity information and better capture facial details. We also improve the loss function to alleviate the interference of the background color. Several experiments show that our method can produce better samples with smoother lines that look more like the corresponding person, outperforming state-of-the-art methods both qualitatively and quantitatively. Our method can also be generalized to other styles such as sketch.
Yexiang Liu, Jin Liu 0040, Jie Cao 0002, Junxian Duan, Ran He 0001
FG5
2024 Semantic-Aware Detail Enhancement for Blind Face Restoration
abstract
The goal of Blind Face Restoration is to recover high-quality images from low-quality images suffering from unknown degradations, posing a significantly challenging problem. In recent years, numerous BFR methods have been proposed, achieving significant success. However, faces possess a unique facial topology, and subtle differences in texture, slight structural imbalances, and minimal asymmetry are easily perceptible in the restored face images. Previous methods often struggle to generate realistically high-quality images from real-world low-quality images and fail to preserve fine features. To more effectively restore image details and textures, providing a more natural and realistic restoration effect, we integrate facial semantic information as prior knowledge into the blind face restoration task. We employ a multi-head cross-attention mechanism to simultaneously consider facial semantic information and context information for modeling. Additionally, we introduce a local detail enhancement module specifically designed to enhance the processing capability of details around the eyes and mouth. Experimental results indicate that our proposed method recovers facial images on synthetic and real datasets more realistically and with higher fidelity.
Xiaoqiang Zhou, Jie Cao 0002, Huaibo Huang, Aihua Zheng, Ran He 0001
FG6
2024 Parallel Augmentation and Dual Enhancement for Occluded Person Re-Identification
abstract
Occluded person re-identification (Re-ID), the task of searching for the same person’s images in occluded environments, has attracted lots of attention in the past decades. Recent approaches concentrate on improving performance on occluded data by data/feature augmentation or using extra models to predict occlusions. However, they ignore the imbalance problem in this task and can not fully utilize the information from the training data. To alleviate these two issues, we propose a simple yet effective method with Parallel Augmentation and Dual Enhancement (PADE), which is robust on both occluded and non-occluded data and does not require any auxiliary clues. First, we design a parallel augmentation mechanism (PAM) to generate more suitable occluded data to mitigate the negative effects of unbalanced data. Second, we propose the global and local dual enhancement strategy (DES) to promote the context information and details. Experimental results on three widely used occluded datasets and two non-occluded datasets validate the effectiveness of our method. The code is available at PADE (GitHub).
Zi Wang 0013, Huaibo Huang, Aihua Zheng, Chenglong Li 0002, Ran He 0001
ICASSP5
2024 ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion Models
abstract
In this work, we investigate the capability of generating images from pre-trained diffusion models at much higher resolutions than the training image sizes. In addition, the generated images should have arbitrary image aspect ratios. When generating images directly at a higher resolution, 1024 x 1024, with the pre-trained Stable Diffusion using training images of resolution 512 x 512, we observe persistent problems of object repetition and unreasonable object structures. Existing works for higher-resolution generation, such as attention-based and joint-diffusion approaches, cannot well address these issues. As a new perspective, we examine the structural components of the U-Net in diffusion models and identify the crucial cause as the limited perception field of convolutional kernels. Based on this key observation, we propose a simple yet effective re-dilation that can dynamically adjust the convolutional perception field during inference. We further propose the dispersed convolution and noise-damped classifier-free guidance, which can enable ultra-high-resolution image generation (e.g., 4096 x 4096). Notably, our approach does not require any training or optimization. Extensive experiments demonstrate that our approach can address the repetition issue well and achieve state-of-the-art performance on higher-resolution image synthesis, especially in texture details. Our work also suggests that a pre-trained diffusion model trained on low-resolution images can be directly used for high-resolution visual generation without further tuning, which may provide insights for future research on ultra-high-resolution image and video synthesis. More results are available at the anonymous website: https://scalecrafter.github.io/ScaleCrafter/
Yingqing He, Shaoshu Yang, Haoxin Chen, Xiaodong Cun, Menghan Xia, Yong Zhang 0034, Xintao Wang 0002, Ran He 0001, Qifeng Chen 0001, Ying Shan
ICLR8
2024 Towards Eliminating Hard Label Constraints in Gradient Inversion Attacks
abstract
Gradient inversion attacks aim to reconstruct local training data from intermediate gradients exposed in the federated learning framework. Despite successful attacks, all previous methods, starting from reconstructing a single data point and then relaxing the single-image limit to batch level, are only tested under hard label constraints. Even for single-image reconstruction, we still lack an analysis-based algorithm to recover augmented soft labels. In this work, we change the focus from enlarging batchsize to investigating the hard label constraints, considering a more realistic circumstance where label smoothing and mixup techniques are used in the training process. In particular, we are the first to initiate a novel algorithm to simultaneously recover the ground-truth augmented label and the input feature of the last fully-connected layer from single-input gradients, and provide a necessary condition for any analytical-based label recovery methods. Extensive experiments testify to the label recovery accuracy, as well as the benefits to the following image reconstruction. We believe soft labels in classification tasks are worth further attention in gradient inversion attacks.
Yanbo Wang 0004, Jian Liang 0001, Ran He 0001
ICLR3
2024 A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
abstract
Contrastive Language-Image Pretraining (CLIP) has gained popularity for its remarkable zero-shot capacity. Recent research has focused on developing efficient fine-tuning methods, such as prompt learning and adapter, to enhance CLIP's performance in downstream tasks. However, these methods still require additional training time and computational resources, which is undesirable for devices with limited resources. In this paper, we revisit a classical algorithm, Gaussian Discriminant Analysis (GDA), and apply it to the downstream classification of CLIP. Typically, GDA assumes that features of each class follow Gaussian distributions with identical covariance. By leveraging Bayes' formula, the classifier can be expressed in terms of the class means and covariance, which can be estimated from the data without the need for training. To integrate knowledge from both visual and textual modalities, we ensemble it with the original zero-shot classifier within CLIP. Extensive results on 17 datasets validate that our method surpasses or achieves comparable results with state-of-the-art methods on few-shot classification, imbalanced learning, and out-of-distribution generalization. In addition, we extend our method to base-to-new generalization and unsupervised learning, once again demonstrating its superiority over competing approaches. Our code is publicly available at https://github.com/mrflogs/ICLR24.
Zhengbo Wang, Jian Liang 0001, Lijun Sheng, Ran He 0001, Zilei Wang, Tieniu Tan
ICLR4
2024 Thought Propagation: an Analogical Approach to Complex Reasoning with Large Language Models
abstract
Large Language Models (LLMs) have achieved remarkable success in reasoning tasks with the development of prompting methods. However, existing prompting approaches cannot reuse insights of solving similar problems and suffer from accumulated errors in multi-step reasoning, since they prompt LLMs to reason \textit{from scratch}. To address these issues, we propose \textbf{\textit{Thought Propagation} (TP)}, which explores the analogous problems and leverages their solutions to enhance the complex reasoning ability of LLMs. These analogous problems are related to the input one, with reusable solutions and problem-solving strategies. Thus, it is promising to propagate insights of solving previous analogous problems to inspire new problem-solving. To achieve this, TP first prompts LLMs to propose and solve a set of analogous problems that are related to the input one. Then, TP reuses the results of analogous problems to directly yield a new solution or derive a knowledge-intensive plan for execution to amend the initial solution obtained from scratch. TP is compatible with existing prompting approaches, allowing plug-and-play generalization and enhancement in a wide range of tasks without much labor in task-specific prompt engineering. Experiments across three challenging tasks demonstrate TP enjoys a substantial improvement over the baselines by an average of 12\% absolute increase in finding the optimal solutions in Shortest-path Reasoning, 13\% improvement of human preference in Creative Writing, and 15\% enhancement in the task completion rate of LLM-Agent Planning.
Junchi Yu, Ran He 0001, Rex Ying
ICLR2
2024 Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
abstract
The emergence of vision-language models, such as CLIP, has spurred a significant research effort towards their application for downstream supervised learning tasks. Although some previous studies have explored the unsupervised fine-tuning of CLIP, they often rely on prior knowledge in the form of class names associated with ground truth labels. This paper explores a realistic unsupervised fine-tuning scenario, considering the presence of out-of-distribution samples from unknown classes within the unlabeled data. In particular, we focus on simultaneously enhancing out-of-distribution detection and the recognition of instances associated with known classes. To tackle this problem, we present a simple, efficient, and effective approach called Universal Entropy Optimization (UEO). UEO leverages sample-level confidence to approximately minimize the conditional entropy of confident instances and maximize the marginal entropy of less confident instances. Apart from optimizing the textual prompt, UEO incorporates optimization of channel-wise affine transformations within the visual branch of CLIP. Extensive experiments across 15 domains and 4 different types of prior knowledge validate the effectiveness of UEO compared to baseline methods. The code is at https://github.com/tim-learn/UEO.
Jian Liang 0001, Lijun Sheng, Zhengbo Wang, Ran He 0001, Tieniu Tan
ICML4
2024 Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
abstract
With the emergence of pretrained vision-language models (VLMs), considerable efforts have been devoted to fine-tuning them for downstream tasks. Despite the progress made in designing efficient fine-tuning methods, such methods require access to the model’s parameters, which can be challenging as model owners often opt to provide their models as a black box to safeguard model ownership. This paper proposes a Collaborative Fine-Tuning (CraFT) approach for fine-tuning black-box VLMs to downstream tasks, where one only has access to the input prompts and the output predictions of the model. CraFT comprises two modules, a prompt generation module for learning text prompts and a prediction refinement module for enhancing output predictions in residual style. Additionally, we introduce an auxiliary prediction-consistent loss to promote consistent optimization across these modules. These modules are optimized by a novel collaborative training algorithm. Extensive experiments on few-shot classification over 15 datasets demonstrate the superiority of CraFT. The results show that CraFT achieves a decent gain of about 12% with 16-shot datasets and only 8,000 queries. Moreover, CraFT trains faster and uses only about 1/80 of the memory footprint for deployment, while sacrificing only 1.62% compared to the white-box method. Our code is publicly available at https://github.com/mrflogs/CraFT.
Zhengbo Wang, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
ICML3
2024 ZePo: Zero-Shot Portrait Stylization with Faster Sampling
abstract
Diffusion-based text-to-image generation models have significantly advanced the field of art content synthesis. However, current portrait stylization methods generally require either model fine-tuning based on examples or the employment of DDIM Inversion to revert images to noise space, both of which substantially decelerate the image generation process. To overcome these limitations, this paper presents an inversion-free portrait stylization framework based on diffusion models that accomplishes content and style feature fusion in merely four sampling steps. We observed that Latent Consistency Models employing consistency distillation can effectively extract representative Consistency Features from noisy images. To blend the Consistency Features extracted from both content and style images, we introduce a Style Enhancement Attention Control technique that meticulously merges content and style features within the attention space of the target image. Moreover, we propose a feature merging strategy to amalgamate redundant features in Consistency Features, thereby reducing the computational load of attention control. Extensive experiments have validated the effectiveness of our proposed framework in enhancing stylization efficiency and fidelity. The code is available at \url{https://github.com/liujin112/ZePo}.
Jin Liu 0040, Huaibo Huang, Jie Cao 0002, Ran He 0001
ACM Multimedia4
2024 Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
abstract
In the realm of Multimodal Large Language Models (MLLMs), vision-language connector plays a crucial role to link the pre-trained vision encoders with Large Language Models (LLMs). Despite its importance, the vision-language connector has been relatively less explored. In this study, we aim to propose a strong vision-language connector that enables MLLM to simultaneously achieve high accuracy and low computation cost. We first reveal the existence of the visual anchors in Vision Transformer and propose a cost-effective search algorithm to progressively extract them. Building on these findings, we introduce the Anchor Former (AcFormer), a novel vision-language connector designed to leverage the rich prior knowledge obtained from these visual anchors during pretraining, guiding the aggregation of information. Through extensive experimentation, we demonstrate that the proposed method significantly reduces computational costs by nearly two-thirds, while simultaneously outperforming baseline methods. This highlights the effectiveness and efficiency of AcFormer.
Haogeng Liu, Quanzeng You, Yongfei Liu, Huaibo Huang, Ran He 0001, Hongxia Yang
NeurIPS6
2024 Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation
abstract
Recent advancements in 3D content generation have been significant, primarily due to the visual priors provided by pretrained diffusion models. However, large 2D visual models exhibit spatial perception hallucinations, leading to multi-view inconsistency in 3D content generated through Score Distillation Sampling (SDS). This phenomenon, characterized by overfitting to specific views, is referred to as the "Janus Problem". In this work, we investigate the hallucination issues of pretrained models and find that large multimodal models without geometric constraints possess the capability to infer geometric structures, which can be utilized to mitigate multi-view inconsistency. Building on this, we propose a novel tuning-free method. We represent the multimodal inconsistency query information to detect specific hallucinations in 3D content, using this as an enhanced prompt to re-consist the 2D renderings of 3D and jointly optimize the structure and appearance across different views. Our approach does not require 3D training data and can be implemented plug-and-play within existing frameworks. Extensive experiments demonstrate that our method significantly improves the consistency of 3D content generation and specifically mitigates hallucinations caused by pretrained large models, achieving state-of-the-art performance compared to other optimization methods.
Jie Cao 0002, Jin Liu 0040, Xiaoqiang Zhou, Huaibo Huang, Ran He 0001
NeurIPS6
2024 Recognizing Predictive Substructures With Subgraph Information Bottleneck
abstract
The emergence of Graph Convolutional Network (GCN) has greatly boosted the progress of graph learning. However, two disturbing factors, noise and redundancy in graph data, and lack of interpretation for prediction results, impede further development of GCN. One solution is to recognize a predictive yet compressed subgraph to get rid of the noise and redundancy and obtain the interpretable part of the graph. This setting of subgraph is similar to the information bottleneck (IB) principle, which is less studied on graph-structured data and GCN. Inspired by the IB principle, we propose a novel subgraph information bottleneck (SIB) framework to recognize such subgraphs, named IB-subgraph. However, the intractability of mutual information and the discrete nature of graph data makes the objective of SIB notoriously hard to optimize. To this end, we introduce a bilevel optimization scheme coupled with a mutual information estimator for irregular graphs. Moreover, we propose a continuous relaxation for subgraph selection with a connectivity loss for stabilization. We further theoretically prove the error bound of our estimation scheme for mutual information and the noise-invariant nature of IB-subgraph. Extensive experiments on graph learning and large-scale point cloud tasks demonstrate the superior property of IB-subgraph.
Junchi Yu, Tingyang Xu, Yu Rong 0001, Yatao Bian, Junzhou Huang, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 MAPS: A Noise-Robust Progressive Learning Approach for Source-Free Domain Adaptive Keypoint Detection
abstract
Existing cross-domain keypoint detection methods always require accessing the source data during adaptation, which may violate the data privacy law and pose serious security concerns. Instead, this paper considers a realistic problem setting called source-free domain adaptive keypoint detection, where only the well-trained source model is provided to the target domain. For the challenging problem, we first construct a teacher-student learning baseline by stabilizing the predictions under data augmentation and network ensembles. Built on this, we further propose a unified approach, Mixup Augmentation and Progressive Selection (MAPS), to fully exploit the noisy pseudo labels of unlabeled target data during training. On the one hand, MAPS regularizes the model to favor simple linear behavior in-between the target samples via self-mixup augmentation, preventing the model from over-fitting to noisy predictions. On the other hand, MAPS employs the self-paced learning paradigm and progressively selects pseudo-labeled samples from ‘easy’ to ‘hard’ into the training process to reduce noise accumulation. Results on four keypoint detection datasets show that MAPS outperforms the baseline and achieves comparable or even better results in comparison to previous non-source-free counterparts. The code is available athttps://github.com/YuheD/MAPS.
Yuhe Ding, Jian Liang 0001, Bo Jiang 0002, Aihua Zheng, Ran He 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Dynamic Graph Memory Bank for Video Inpainting
abstract
A major challenge of the video inpainting task is aggregating spatial and temporal information in the corrupted video effectively. In this paper, we propose a dynamic graph memory bank to settle this challenge. To model the long-range temporal dependency, a memory bank is built and updated dynamically with the input visual information flow. The relationships among the memory items are modeled through a graph-based message propagation. Benefiting from the dynamic graph memory bank, both contents and their relationships in the corrupted video are well exploited as the inpainting process going on. Besides, the spatial misalignment across different frames may degrade the quality of features in the dynamic graph memory bank. To alleviate this issue, we propose a motion-guided feature alignment module. The proposed module cooperates with the dynamic graph memory bank to improve the network’s information aggregation ability in spatial and temporal dimensions. Extensive experiments on the YouTube-VOS and DAVIS datasets demonstrate the superiority of our approach when compared with the state-of-the-arts.
Xiaoqiang Zhou, Chaoyou Fu, Huaibo Huang, Ran He 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Attribute-Guided Cross-Modal Interaction and Enhancement for Audio-Visual Matching
abstract
Audio-visual matching is an essential task that measures the correlation between audio clips and visual images. However, current methods rely solely on the joint embedding of global features from audio clips and face image pairs to learn semantic correlations. This approach overlooks the importance of high-confidence correlations and discrepancies of local subtle features, which are crucial for cross-modal matching. To address this issue, we propose a novel Attribute-guided Cross-modal Interaction and Enhancement Network (ACIENet), which employs multiple attributes to explore the associations of different key local subtle features. The ACIENet contains two novel modules: the Attribute-guided Interaction (AGI) module and the Attribute-guided Enhancement (AGE) module. The AGI module employs global feature alignment similarity to guide cross-modal local feature interactions, which enhances cross-modal association features for the same identity and expands cross-modal distinctive features for different identities. Additionally, the interactive features and original features are fused to ensure intra-class discriminability and inter-class correspondence. The AGE module captures subtle attribute-related features by using an attribute-driven network, thereby enhancing discrimination at the attribute level. Specifically, it strengthens the combined attribute-related features of gender and nationality. To prevent interference between multiple attribute features, we design a multi-attribute learning network as a parallel framework. Experiments conducted on a public benchmark dataset demonstrate the efficacy of the ACIENet method in different scenarios. Code and models are available at https://github.com/w1018979952/ACIENet.
Jiaxiang Wang 0001, Aihua Zheng, Yan Yan 0002, Ran He 0001, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.4
2024 Multi-Faceted Knowledge-Driven Graph Neural Network for Iris Segmentation
abstract
Accurate iris segmentation, especially around the iris inner and outer boundaries, is still a formidable challenge. Pixels within these areas are difficult to semantically distinguish since they have similar visual characteristics and close spatial positions. To tackle this problem, the paper proposes an iris segmentation graph neural network (ISeGraph) for accurate segmentation. ISeGraph regards individual pixels as nodes within the graph and constructs self-adaptive edges according to multi-faceted knowledge, including visual similarity, positional correlation, and semantic consistency for feature aggregation. Specifically, visual similarity strengthens the connections between nodes sharing similar visual characteristics, while positional correlation assigns weights according to the spatial distance between nodes. In contrast to the above knowledge, semantic consistency maps nodes into a semantic space and learns pseudo-labels to define relationships based on label consistency. ISeGraph leverages multi-faceted knowledge to generate self-adaptive relationships for accurate iris segmentation. Furthermore, a pixel-wise adaptive normalization module is developed to increase the feature discriminability. It takes informative features in the shallow layer as a reference to improve the segmentation features from a statistical perspective. Experimental results on three iris datasets illustrate that the proposed method achieves superior performance in iris segmentation, increasing the segmentation accuracy in areas near the iris boundaries.
Jianze Wei, Yunlong Wang 0003, Xingyu Gao 0001, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2024 RISTRA: Recursive Image Super-Resolution Transformer With Relativistic Assessment
abstract
Many recent image restoration methods use Transformer as the backbone network and redesign the Transformer blocks. Differently, we explore the parameter-sharing mechanism over Transformer blocks and propose a dynamic recursive process to address the image super-resolution task efficiently. We firstly present a Recursive Image Super-resolution Transformer (RIST). By sharing the weights across different blocks, a plain forward process through the whole Transformer network can be folded into recursive iterations through a Transformer block. Such a parameter-sharing based recursive process can not only reduce the model size greatly, but also enable restoring images progressively. Features in the recursive process are modeled as a sequence and propagated with a temporal attention network. Besides, by analyzing the prediction variation across different iterations in RIST, we design a dynamic recursive process that can allocate adaptive computation costs to different samples. Specifically, a quality assessment network estimates the restoration quality and terminates the recursive process dynamically. We propose a relativistic learning strategy to simplify the objective from absolute image quality assessment to relativistic quality comparison. The proposed Recursive Image Super-resolution Transformer with Relativistic Assessment (RISTRA) reduces the model size greatly with the parameter-sharing mechanism, and achieves an instance-wise dynamic restoration process as well. Extensive experiments on several image super-resolution benchmarks show the superiority of our approach over state-of-the-art counterparts
Xiaoqiang Zhou, Huaibo Huang, Zilei Wang, Ran He 0001
IEEE Trans. Multim.4
2023 Mind the Label Shift of Augmentation-based Graph OOD Generalization
abstract
Out-of-distribution (OOD) generalization is an important issue for Graph Neural Networks (GNNs). Recent works employ different graph editions to generate augmented environments and learn an invariant GNN for generalization. However, the label shift usually occurs in augmentation since graph structural edition inevitably alters the graph label. This brings inconsistent predictive relationships among augmented environments, which is harmful to generalization. To address this issue, we propose LiSA, which generates label-invariant augmentations to facilitate graph OOD generalization. Instead of resorting to graph editions, LiSA exploits Label-invariant Subgraphs of the training graphs to construct Augmented environments. Specifically, LiSA first designs the variational subgraph generators to extract locally predictive patterns and construct multiple label-invariant subgraphs efficiently. Then, the subgraphs produced by different generators are collected to build different augmented environments. To promote diversity among augmented environments, LiSA further introduces a tractable energy-based regularization to enlarge pair-wise distances between the distributions of environments. In this manner, LiSA generates diverse augmented environments with a consistent predictive relationship and facilitates learning an invariant GNN. Extensive experiments on node-level and graph-level OOD benchmarks show that LiSA achieves impressive generalization performance with different GNN backbones. Code is available on https://github.com/Samyu0304/LiSA.
Junchi Yu, Jian Liang 0001, Ran He 0001
CVPR3
2023 Modify: Model-Driven Face Stylization Without Style Images
abstract
Existing face stylization methods always acquire the presence of the target (style) domain during the translation process, which violates privacy regulations and limits their applicability in real-world systems. To address this issue, we propose a new method called MODel-drIven Face stYlization (MODIFY), which relies on the generative model to bypass the dependence of the target images. Briefly, MODIFY first trains a generative model in the target domain and then translates a source input to the target domain via the provided style model. To preserve the multimodal style information, MODIFY further introduces an additional remapping network, mapping a known continuous distribution into the encoder’s embedding space. During translation in the source domain, MODIFY fine-tunes the encoder module within the target style-persevering model to capture the content of the source input as precisely as possible. Our method is extremely simple and satisfies versatile training modes for face stylization. Experimental results on several different datasets validate the effectiveness of MODIFY for unsupervised face stylization. Code will be released at https://github.com/YuheD/MODIFY.
Yuhe Ding, Jian Liang 0001, Jie Cao 0002, Aihua Zheng, Ran He 0001
ICASSP5
2023 Pluralistic Aging Diffusion Autoencoder
abstract
Face aging is an ill-posed problem because multiple plausible aging patterns may correspond to a given input. Most existing methods often produce one deterministic estimation. This paper proposes a novel CLIP-driven Pluralistic Aging Diffusion Autoencoder (PADA) to enhance the diversity of aging patterns. First, we employ diffusion models to generate diverse low-level aging details via a sequential denoising reverse process. Second, we present Probabilistic Aging Embedding (PAE) to capture diverse high-level aging patterns, which represents age information as probabilistic distributions in the common CLIP latent space. A text-guided KL-divergence loss is designed to guide this learning. Our method can achieve pluralistic face aging conditioned on open-world aging texts and arbitrary unseen face images. Qualitative and quantitative experiments demonstrate that our method can generate more diverse and high-quality plausible aging results.
Peipei Li 0002, Rui Wang 0124, Huaibo Huang, Ran He 0001, Zhaofeng He 0001
ICCV4
2023 TALL: Thumbnail Layout for Deepfake Video Detection
abstract
The growing threats of deepfakes to society and cybersecurity have raised enormous public concerns, and increasing efforts have been devoted to this critical topic of deepfake video detection. Existing video methods achieve good performance but are computationally intensive. This paper introduces a simple yet effective strategy named Thumbnail Layout (TALL), which transforms a video clip into a pre-defined layout to realize the preservation of spatial and temporal dependencies. Specifically, consecutive frames are masked in a fixed position in each frame to improve generalization, then resized to sub-images and rearranged into a pre-defined layout as the thumbnail. TALL is model-agnostic and extremely simple by only modifying a few lines of code. Inspired by the success of vision transformers, we incorporate TALL into Swin Transformer, forming an efficient and effective method TALL-Swin. Extensive experiments on intra-dataset and cross-dataset validate the validity and superiority of TALL and SOTA TALL-Swin. TALL-Swin achieves 90.79% AUC on the challenging cross-dataset task, FaceForensics++ → CelebDF. The code is available at https://github.com/rainy-xu/TALL4Deepfake.
Jian Liang 0001, Gengyun Jia, Zimin (Max) Yang, Ran He 0001
ICCV6
2023 Lightweight Vision Transformer with Bidirectional Interaction
abstract
Recent advancements in vision backbones have significantly improved their performance by simultaneously modeling images’ local and global contexts. However, the bidirectional interaction between these two contexts has not been well explored and exploited, which is important in the human visual system. This paper proposes a **F**ully **A**daptive **S**elf-**A**ttention (FASA) mechanism for vision transformer to model the local and global information as well as the bidirectional interaction between them in context-aware ways. Specifically, FASA employs self-modulated convolutions to adaptively extract local representation while utilizing self-attention in down-sampled space to extract global representation. Subsequently, it conducts a bidirectional adaptation process between local and global representation to model their interaction. In addition, we introduce a fine-grained downsampling strategy to enhance the down-sampled self-attention mechanism for finer-grained global perception capability. Based on FASA, we develop a family of lightweight vision backbones, **F**ully **A**daptive **T**ransformer (FAT) family. Extensive experiments on multiple vision tasks demonstrate that FAT achieves impressive performance. Notably, FAT accomplishes a **77.6%** accuracy on ImageNet-1K using only **4.5M** parameters and **0.7G** FLOPs, which surpasses the most advanced ConvNets and Transformers with similar model size and computational costs. Moreover, our model exhibits faster speed on modern GPU compared to other models.
Qihang Fan, Huaibo Huang, Xiaoqiang Zhou, Ran He 0001
NeurIPS4
2023 Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification
abstract
We present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspire us to leverage the rich ordinal priors in human language by converting the original task into a vision-language alignment task. Consequently, we propose L2RCLIP, which fully utilizes the language priors from two perspectives. First, we introduce a complementary prompt tuning technique called RankFormer, designed to enhance the ordering relation of original rank prompts. It employs token-level attention with residual-style prompt blending in the word embedding space. Second, to further incorporate language priors, we revisit the approximate bound optimization of vanilla cross-entropy loss and restructure it within the cross-modal embedding space. Consequently, we propose a cross-modal ordinal pairwise loss to refine the CLIP feature space, where texts and images maintain both semantic alignment and ordering alignment. Extensive experiments on three ordinal classification tasks, including facial age estimation, historical color image (HCI) classification, and aesthetic assessment demonstrate its promising performance.
Rui Wang 0124, Peipei Li 0002, Huaibo Huang, Chunshui Cao, Ran He 0001, Zhaofeng He 0001
NeurIPS5
2023 Diverse features discovery transformer for pedestrian attribute recognition
Aihua Zheng, Jiaxiang Wang 0001, Huaibo Huang, Ran He 0001, Amir Hussain 0001
Eng. Appl. Artif. Intell.5
2023 ProxyMix: Proxy-based Mixup training with label refinery for source-free domain adaptation
Yuhe Ding, Lijun Sheng, Jian Liang 0001, Aihua Zheng, Ran He 0001
Neural Networks5
2023 ScoreMix: A Scalable Augmentation Strategy for Training GANs With Limited Data
abstract
Generative Adversarial Networks (GANs) typically suffer from overfitting when limited training data is available. To facilitate GAN training, current methods propose to use data-specific augmentation techniques. Despite the effectiveness, it is difficult for these methods to scale to practical applications. In this article, we present ScoreMix, a novel and scalable data augmentation approach for various image synthesis tasks. We first produce augmented samples using the convex combinations of the real samples. Then, we optimize the augmented samples by minimizing the norms of the data scores, i.e., the gradients of the log-density functions. This procedure enforces the augmented samples close to the data manifold. To estimate the scores, we train a deep estimation network with multi-scale score matching. For different image synthesis tasks, we train the score estimation network using different data. We do not require the tuning of the hyperparameters or modifications to the network architecture. The ScoreMix method effectively increases the diversity of data and reduces the overfitting problem. Moreover, it can be easily incorporated into existing GAN models with minor modifications. Experimental results on numerous tasks demonstrate that GAN models equipped with the ScoreMix method achieve significant improvements.
Jie Cao 0002, Mandi Luo, Junchi Yu, Ming-Hsuan Yang 0001, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Towards Lightweight Pixel-Wise Hallucination for Heterogeneous Face Recognition
abstract
Cross-spectral face hallucination is an intuitive way to mitigate the modality discrepancy in Heterogeneous Face Recognition (HFR). However, due to imaging differences, the hallucination inevitably suffers from a shape misalignment between paired heterogeneous images. Rather than building complicated architectures to circumvent the problem like previous works, we propose a simple yet effective method called Shape Alignment FacE (SAFE). Specifically, given an image, we align its shape to that of the paired one under the assistance of a 3D face model. The produced aligned pair enables us to train a lightweight generator that solely concentrates on spectrum translation with a pixel-wise supervision. However, since the 3D face model is powerless to attributes like the hair and glasses, there are still pixel discrepancies between the aligned pair. Given that, in the image space, we introduce a probabilistic pixel-wise loss that incorporates the discrepancies into a probabilistic distribution. Moreover, in order to alleviate the influence of the shape misalignment on spectrum translation, a spectrum optimal transport is performed in a shape-irrelevant latent space. Note that, in the final inference phase, except the lightweight generator, all other auxiliary modules are discarded. In addition to superior performance in qualitative synthesis and quantitative recognition, extensive experiments on 6 datasets demonstrate that our method also gains other two distinct advantages over existing state-of-the-art counterparts. The first is using a more lightweight generator. Compared with the state-of-the-art method, our method can achieve higher recognition results with 128x fewer parameters and 63x fewer FLOPs with only 4.58 ms latency on a single TITAN-XP. The second is training on low-shot datasets such as Oulu-CASIA NIR-VIS that just contains 1,920 images from 20 identities. To the best of our knowledge, we are the first that can perform well on such a small-scale dataset. These advantages make our method more practical in the real world and further push boundaries of heterogeneous face recognition.
Chaoyou Fu, Xiaoqiang Zhou, Weizan He, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Memory Uncertainty Learning for Real-World Single Image Deraining
abstract
Single image deraining has witnessed dramatic improvements by training deep neural networks on large-scale synthetic data. However, due to the discrepancy between authentic and synthetic rain images, it is challenging to directly extend existing methods to real-world scenes. To address this issue, we propose a memory-uncertainty guided semi-supervised method to learn rain properties simultaneously from synthetic and real data. The key aspect is developing a stochastic memory network that is equipped with memory modules to record prototypical rain patterns. The memory modules are updated in a self-supervised way, allowing the network to comprehensively capture rainy styles without the need for clean labels. The memory items are read stochastically according to their similarities with rain representations, leading to diverse predictions and efficient uncertainty estimation. Furthermore, we present an uncertainty-aware self-training mechanism to transfer knowledge from supervised deraining to unsupervised cases. An additional target network is adopted to produce pseudo-labels for unlabeled data, of which the incorrect ones are rectified by uncertainty estimates. Finally, we construct a new large-scale image deraining dataset of 10.2 k real rain images, significantly improving the diversity of real rain scenes. Experiments show that our method achieves more appealing results for real-world rain removal than recent state-of-the-art methods.
Huaibo Huang, Mandi Luo, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Iterative embedding distillation for open world vehicle recognition
Junxian Duan, Xiang Wu 0001, Yibo Hu 0001, Chaoyou Fu, Zi Wang 0013, Ran He 0001
Pattern Recognit.6
2023 Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric Synthesis
abstract
Existing automated dubbing methods are usually designed for Professionally Generated Content (PGC) production, which requires massive training data and training time to learn a person-specific audio-video mapping. In this paper, we investigate an audio-driven dubbing method that is more feasible for User Generated Content (UGC) production. There are two unique challenges to design a method for UGC: 1) the appearances of speakers are diverse and arbitrary as the method needs to generalize across users; 2) the available video data of one speaker are very limited. In order to tackle the above challenges, we first introduce a new Style Translation Network to integrate the speaking style of the target and the speaking content of the source via a cross-modal AdaIN module. It enables our model to quickly adapt to a new speaker. Then, we further develop a semi-parametric video renderer, which takes full advantage of the limited training data of the unseen speaker via a video-level retrieve-warp-refine pipeline. Finally, we propose a temporal regularization for the semi-parametric renderer, generating more continuous videos. Extensive experiments show that our method generates videos that accurately preserve various speaking styles, yet with considerably lower amount of training data and training time in comparison to existing methods. Besides, our method achieves a faster testing speed than most recent methods.
Linsen Song, Wayne Wu, Chaoyou Fu, Chen Change Loy, Ran He 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Contextual Measures for Iris Recognition
abstract
The iris patterns of the human contain a large amount of randomly distributed and irregularly shaped microstructures. These microstructures make the human iris informative biometric traits. To learn identity representation from them, this paper regards each iris region as a potential microstructure and proposes contextual measures (CM) to model the correlations between them. CM adopts two parallel branches to learn global and local contexts in iris image. The first one is the globally contextual measure branch. It measures the global context involving the relationships between all regions for feature aggregation and is robust to local occlusions. Besides, we improve its spatial perception considering the positional randomness of the microstructures. The other one is the locally contextual measure branch. This branch considers the role of local details in the phenotypic distinctiveness of iris patterns and learns a series of relationship atoms to capture contextual information from a local perspective. In addition, we develop the perturbation bottleneck to make sure that the two branches learn divergent contexts. It introduces perturbation to limit the information flow from input images to identity features, forcing CM to learn discriminative contextual information for iris recognition. Experimental results suggest that global and local contexts are two different clues critical for accurate iris recognition. The superior performance on four benchmark iris datasets demonstrates the effectiveness of the proposed approach in within-database and cross-database scenarios.
Jianze Wei, Yunlong Wang 0003, Huaibo Huang, Ran He 0001, Zhenan Sun, Xingyu Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2023 Masked Relation Learning for DeepFake Detection
abstract
DeepFake detection aims to differentiate falsified faces from real ones. Most approaches formulate it as a binary classification problem by solely mining the local artifacts and inconsistencies of face forgery, which neglect the relation across local regions. Although several recent works explore local relation learning for DeepFake detection, they overlook the propagation of relational information and lead to limited performance gains. To address these issues, this paper provides a new perspective by formulating DeepFake detection as a graph classification problem, in which each facial region corresponds to a vertex. But relational information with large redundancy hinders the expressiveness of graphs. Inspired by the success of masked modeling, we propose Masked Relation Learning which decreases the redundancy to learn informative relational features. Specifically, a spatiotemporal attention module is exploited to learn the attention features of multiple facial regions. A relation learning module masks partial correlations between regions to reduce redundancy and then propagates the relational information across regions to capture the irregularity from a global view of the graph. We empirically discover that a moderate masking rate (e.g., 50%) brings the best performance gain. Experiments verify the effectiveness of Masked Relation Learning and demonstrate that our approach outperforms the state of the art by 2% AUC on the cross-dataset DeepFake video detection. Code will be available athttps://github.com/zimyang/MaskRelation.
Zimin (Max) Yang, Jian Liang 0001, Xiaoyu Zhang 0002, Ran He 0001
IEEE Trans. Inf. Forensics Secur.5
2023 Theme-Aware Aesthetic Distribution Prediction With Full-Resolution Photographs
abstract
Aesthetic quality assessment (AQA) is a challenging task due to complex aesthetic factors. Currently, it is common to conduct AQA using deep neural networks (DNNs) that require fixed-size inputs. The existing methods mainly transform images by resizing, cropping, and padding or use adaptive pooling to alternately capture the aesthetic features from fixed-size inputs. However, these transformations potentially damage aesthetic features. To address this issue, we propose a simple but effective method to accomplish full-resolution image AQA by combining image padding with region of image (RoM) pooling. Padding turns inputs into the same size. RoM pooling pools image features and discards extra padded features to eliminate the side effects of padding. In addition, the image aspect ratios are encoded and fused with visual features to remedy the shape information loss of RoM pooling. Furthermore, we observe that the same image may receive different aesthetic evaluations under different themes, which we call the theme criterion bias. Hence, a theme-aware model that uses theme information to guide model predictions is proposed. Finally, we design an attention-based feature fusion module to effectively use both the shape and theme information. Extensive experiments prove the effectiveness of the proposed method over state-of-the-art methods.
Gengyun Jia, Peipei Li 0002, Ran He 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identification
abstract
Multi-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations for multi-modal person Re-ID: Interact, Embed, and EnlargE (IEEE). First, we propose a cross-modal interacting module to exchange useful information between different modalities in the feature extraction phase. Second, we propose a relation-based embedding module to enhance the richness of feature descriptors by embedding the global feature into the fine-grained local information. Finally, we propose multi-modal margin loss to force the network to learn modality-specific information for each modality by enlarging the intra-class discrepancy. Superior performance on multi-modal Re-ID dataset RGBNT201 and three constructed Re-ID datasets validate the effectiveness of the proposed method compared with the state-of-the-art approaches.
Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Ran He 0001, Jin Tang 0001
AAAI4
2022 Confidence-Calibrated Face Image Forgery Detection with Contrastive Representation Distillation
Puning Yang, Huaibo Huang, Zhiyong Wang 0001, Aijing Yu, Ran He 0001
ACCV (4)5
2022 Few-shot Backdoor Defense Using Shapley Estimation
abstract
Deep neural networks have achieved impressive performance in a variety of tasks over the last decade, such as autonomous driving, face recognition, and medical diagnosis. However, prior works show that deep neural networks are easily manipulated into specific, attacker-decided behaviors in the inference stage by backdoor attacks which inject malicious small hidden triggers into model training, raising serious security threats. To determine the triggered neurons and protect against backdoor attacks, we exploit Shapley value and develop a new approach called Shapley Pruning (ShapPruning) that successfully mitigates backdoor attacks from models in a data-insufficient situation (1 image per class or even free of data). Considering the interaction between neurons, ShapPruning identifies the few infected neurons (under 1 % of all neurons) and manages to protect the model's structure and accuracy after pruning as many infected neurons as possible. To accelerate ShapPruning, we further propose discarding threshold and ∊ -greedy strategy to accelerate Shapley estimation, making it possible to repair poisoned models with only several minutes. Experiments demonstrate the effectiveness and robustness of our method against various attacks and tasks compared to existing methods.
Jiyang Guan, Zhuozhuo Tu, Ran He 0001, Dacheng Tao
CVPR3
2022 Rethinking Image Cropping: Exploring Diverse Compositions from Global Views
abstract
Existing image cropping works mainly use anchor evaluation methods or coordinate regression methods. However, it is difficult for pre-defined anchors to cover good crops globally, and the regression methods ignore the cropping diversity. In this paper, we regard image cropping as a set prediction problem. A set of crops regressed from multiple learnable anchors is matched with the labeled good crops, and a classifier is trained using the matching results to select a valid subset from all the predictions. This new perspective equips our model with globality and diversity, mitigating the shortcomings but inherit the strengthens of previous methods. Despite the advantages, the set prediction method causes inconsistency between the validity labels and the crops. To deal with this problem, we propose to smooth the validity labels with two different methods. The first method that uses crop qualities as direct guidance is designed for the datasets with nearly dense quality labels. The second method based on the self distillation can be used in sparsely labeled datasets. Experimental results on the public datasets show the merits of our approach over state-of-the-art counterparts.
Gengyun Jia, Huaibo Huang, Chaoyou Fu, Ran He 0001
CVPR4
2022 DINE: Domain Adaptation from Single and Multiple Black-box Predictors
abstract
To ease the burden of labeling, unsupervised domain adaptation (UDA) aims to transfer knowledge in previous and related labeled datasets (sources) to a new unlabeled dataset (target). Despite impressive progress, prior methods always need to access the raw source data and develop data-dependent alignment approaches to recognize the target samples in a transductive learning manner, which may raise privacy concerns from source individuals. Several recent studies resort to an alternative solution by exploiting the well-trained white-box model from the source domain, yet, it may still leak the raw data via generative adversarial learning. This paper studies a practical and interesting setting for UDA, where only black-box source models (i.e., only network predictions are available) are provided during adaptation in the target domain. To solve this problem, we propose a new two-step knowledge adaptation framework called DIstill and fine-tuNE (DINE). Taking into consideration the target data structure, DINE first distills the knowledge from the source predictor to a customized target model, then fine-tunes the distilled model to further fit the target domain. Besides, neural networks are not required to be identical across domains in DINE, even allowing effective adaptation on a low-resource device. Empirical results on three UDA scenarios (i.e., single-source, multisource, and partial-set) confirm that DINE achieves highly competitive performance compared to state-of-the-art data-dependent approaches. Code is available at https://github.com/tim-learn/DINE/.
Jian Liang 0001, Dapeng Hu, Jiashi Feng, Ran He 0001
CVPR4
2022 Improving Subgraph Recognition with Variational Graph Information Bottleneck
abstract
Subgraph recognition aims at discovering a compressed substructure of a graph that is most informative to the graph property. It can be formulated by optimizing Graph Information Bottleneck (GIB) with a mutual information estimator. However, GIB suffers from training instability and degenerated results due to its intrinsic optimization process. To tackle these issues, we reformulate the subgraph recognition problem into two steps: graph perturbation and subgraph selection, leading to a novel Variational Graph Information Bottleneck (VGIB) framework. VGIB first employs the noise injection to modulate the information flow from the input graph to the perturbed graph. Then, the perturbed graph is encouraged to be informative to the graph property. VGIB further obtains the desired subgraph by filtering out the noise in the perturbed graph. With the customized noise prior for each input, the VGIB objective is endowed with a tractable variational upper bound, leading to a superior empirical performance as well as theoretical properties. Extensive experiments on graph interpretation, explainability of Graph Neural Networks, and graph classification show that VGIB finds better subgraphs than existing methods11Code is avaliable on https://github.com/Samyu0304/VGIB.
Junchi Yu, Jie Cao 0002, Ran He 0001
CVPR3
2022 Fine-Grained Cross-Modal Retrieval with Triple-Streamed Memory Fusion Transformer Encoder
abstract
Recently, the powerful attention mechanism has been wildly used to learn the fine-grained cross-modal correspondences. However, the trade-off between effectiveness and efficiency sometimes bothers existing attention-mechanism-based meth-ods. To address this deficiency, we propose a novel Triple-streamed architecture with a newly designed Memory fusion Transformer Encoder (Tri-MTE) for fine-grained cross-modal retrieval. Specifically, the whole model reserves the “late fusion” strategy thus ensuring efficiency. To strengthen the inter-modality interaction and improve the effectiveness, a memory fusion stream is designed and inserted between the modality streams to remember the modality-irrelevant infor-mation. Encoding such information to the modality represen-tation would significantly enhance the cross-modal retrieval performance. Finally, a bionic memory activation constrain-t is proposed to aid the learning procedure. Extensive ex-periments on two benchmark datasets show that the proposed method achieves promising results.
Weikuo Guo, Huaibo Huang, Xiangwei Kong 0001, Ran He 0001
ICME4
2022 Order-aware Human Interaction Manipulation
abstract
The majority of current techniques for pose transfer disregard the interactions between the transferred person and the surrounding instances, resulting in context inconsistency when applied to complicated situations. To tackle this issue, we propose InterOrderNet, a novel framework to perform order-aware interaction learning. The proposed InterOrderNet learns the relative order on the direction of the z-axis among instances to describe instance-level occlusions. Not only does learning this order guarantee the context consistency of human pose transfer, but it also enhances its generalization to natural scenes. Additionally, we present a novel unsupervised method, named Imitative Contrastive Learning, which sidesteps the requirements of order annotations. Existing pose transfer methods are easy to be integrated into the proposed InterOrderNet. Extensive experiments demonstrate that InterOrderNet enables these methods to perform interaction manipulation.
Mandi Luo, Jie Cao 0002, Ran He 0001
ACM Multimedia3
2022 Are You Stealing My Model? Sample Correlation for Fingerprinting Deep Neural Networks
abstract
An off-the-shelf model as a commercial service could be stolen by model stealing attacks, posing great threats to the rights of the model owner. Model fingerprinting aims to verify whether a suspect model is stolen from the victim model, which gains more and more attention nowadays. Previous methods always leverage the transferable adversarial examples as the model fingerprint, which is sensitive to adversarial defense or transfer learning scenarios. To address this issue, we consider the pairwise relationship between samples instead and propose a novel yet simple model stealing detection method based on SAmple Correlation (SAC). Specifically, we present SAC-w that selects wrongly classified normal samples as model inputs and calculates the mean correlation among their model outputs. To reduce the training time, we further develop SAC-m that selects CutMix Augmented samples as model inputs, without the need for training the surrogate models or generating adversarial examples. Extensive results validate that SAC successfully defends against various model stealing attacks, even including adversarial training or transfer learning, and detects the stolen models with the best performance in terms of AUC across different datasets and model architectures. The codes are available at https://github.com/guanjiyang/SAC.
Jiyang Guan, Jian Liang 0001, Ran He 0001
NeurIPS3
2022 Orthogonal Transformer: An Efficient Vision Transformer Backbone with Token Orthogonalization
abstract
We present a general vision transformer backbone, called as Orthogonal Transformer, in pursuit of both efficiency and effectiveness. A major challenge for vision transformer is that self-attention, as the key element in capturing long-range dependency, is very computationally expensive for dense prediction tasks (e.g., object detection). Coarse global self-attention and local self-attention are then designed to reduce the cost, but they suffer from either neglecting local correlations or hurting global modeling. We present an orthogonal self-attention mechanism to alleviate these issues. Specifically, self-attention is computed in the orthogonal space that is reversible to the spatial domain but has much lower resolution. The capabilities of learning global dependency and exploring local correlations are maintained because every orthogonal token in self-attention can attend to the entire visual tokens. Remarkably, orthogonality is realized by constructing an endogenously orthogonal matrix that is friendly to neural networks and can be optimized as arbitrary orthogonal matrices. We also introduce Positional MLP to incorporate position information for arbitrary input resolutions as well as enhance the capacity of MLP. Finally, we develop a hierarchical architecture for Orthogonal Transformer. Extensive experiments demonstrate its strong performance on a broad range of vision tasks, including image classification, object detection, instance segmentation and semantic segmentation.
Huaibo Huang, Xiaoqiang Zhou, Ran He 0001
NeurIPS3
2022 Style-Based Attentive Network for Real-World Face Hallucination
Mandi Luo, Xin Ma 0031, Huaibo Huang, Ran He 0001
PRCV (4)4
2022 Prior-Guided Multi-scale Fusion Transformer for Face Attribute Recognition
Shaoheng Song, Huaibo Huang, Jiaxiang Wang 0001, Aihua Zheng, Ran He 0001
PRCV (1)5
2022 DVG-Face: Dual Variational Generation for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) refers to matching cross-domain faces and plays a crucial role in public security. Nevertheless, HFR is confronted with challenges from large domain discrepancy and insufficient heterogeneous data. In this paper, we formulate HFR as a dual generation problem, and tackle it via a novel dual variational generation (DVG-Face) framework. Specifically, a dual variational generator is elaborately designed to learn the joint distribution of paired heterogeneous images. However, the small-scale paired heterogeneous training data may limit the identity diversity of sampling. In order to break through the limitation, we propose to integrate abundant identity information of large-scale visible data into the joint distribution. Furthermore, a pairwise identity preserving loss is imposed on the generated paired heterogeneous images to ensure their identity consistency. As a consequence, massive new diverse paired heterogeneous images with the same identity can be generated from noises. The identity consistency and identity diversity properties allow us to employ these generated images to train the HFR network via a contrastive learning mechanism, yielding both domain-invariant and discriminative embedding features. Concretely, the generated paired heterogeneous images are regarded as positive pairs, and the images obtained from different samplings are considered as negative pairs. Our method achieves superior performances over state-of-the-art methods on seven challenging databases belonging to five HFR tasks, including NIR-VIS, Sketch-Photo, Profile-Frontal Photo, Thermal-VIS, and ID-Camera.
Chaoyou Fu, Xiang Wu 0001, Yibo Hu 0001, Huaibo Huang, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Source Data-Absent Unsupervised Domain Adaptation Through Hypothesis Transfer and Labeling Transfer
abstract
Unsupervised domain adaptation (UDA) aims to transfer knowledge from a related but different well-labeled source domain to a new unlabeled target domain. Most existing UDA methods require access to the source data, and thus are not applicable when the data are confidential and not shareable due to privacy concerns. This paper aims to tackle a realistic setting with only a classification model available trained over, instead of accessing to, the source data. To effectively utilize the source model for adaptation, we propose a novel approach called Source HypOthesis Transfer (SHOT), which learns the feature extraction module for the target domain by fitting the target data features to the frozen source classification module (representing classification hypothesis). Specifically, SHOT exploits both information maximization and self-supervised learning for the feature extraction module learning to ensure the target features are implicitly aligned with the features of unseen source data via the same hypothesis. Furthermore, we propose a new labeling transfer strategy, which separates the target data into two splits based on the confidence of predictions (labeling information), and then employ semi-supervised learning to improve the accuracy of less-confident predictions in the target domain. We denote labeling transfer as SHOT++ if the predictions are obtained by SHOT. Extensive experiments on both digit classification and object recognition tasks show that SHOT and SHOT++ achieve results surpassing or comparable to the state-of-the-arts, demonstrating the effectiveness of our approaches for various visual domain adaptation problems. Code will be available at https://github.com/tim-learn/SHOT-plus.
Jian Liang 0001, Dapeng Hu, Yunbo Wang, Ran He 0001, Jiashi Feng
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 PSGAN++: Robust Detail-Preserving Makeup Transfer and Removal
abstract
In this paper, we address the makeup transfer and removal tasks simultaneously, which aim to transfer the makeup from a reference image to a source image and remove the makeup from the with-makeup image respectively. Existing methods have achieved much advancement in constrained scenarios, but it is still very challenging for them to transfer makeup between images with large pose and expression differences, or handle makeup details like blush on cheeks or highlight on the nose. In addition, they are hardly able to control the degree of makeup during transferring or to transfer a specified part in the input face. These defects limit the application of previous makeup transfer methods to real-world scenarios. In this work, we propose a Pose and expression robust Spatial-aware GAN (abbreviated as PSGAN++). PSGAN++ is capable of performing both detail-preserving makeup transfer and effective makeup removal. For makeup transfer, PSGAN++ uses a Makeup Distill Network (MDNet) to extract makeup information, which is embedded into spatial-aware makeup matrices. We also devise an Attentive Makeup Morphing (AMM) module that specifies how the makeup in the source image is morphed from the reference image, and a makeup detail loss to supervise the model within the selected makeup detail area. On the other hand, for makeup removal, PSGAN++ applies an Identity Distill Network (IDNet) to embed the identity information from with-makeup images into identity matrices. Finally, the obtained makeup/identity matrices are fed to a Style Transfer Network (STNet) that is able to edit the feature maps to achieve makeup transfer or removal. To evaluate the effectiveness of our PSGAN++, we collect a Makeup Transfer In the Wild (MT-Wild) dataset that contains images with diverse poses and expressions and a Makeup Transfer High-Resolution (MT-HR) dataset that contains high-resolution images. Experiments demonstrate that PSGAN++ not only achieves state-of-the-art results with fine makeup details even in cases of large pose/expression differences but also can perform partial or degree-controllable makeup transfer. Both the code and the newly collected datasets will be released at https://github.com/wtjiang98/PSGAN.
Si Liu 0001, Chen Gao 0005, Ran He 0001, Jiashi Feng, Bo Li 0006, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Deep momentum uncertainty hashing
Chaoyou Fu, Guoli Wang 0004, Xiang Wu 0001, Qian Zhang 0009, Ran He 0001
Pattern Recognit.5
2022 Structure-aware conditional variational auto-encoder for constrained molecule optimization
Junchi Yu, Tingyang Xu, Yu Rong 0001, Junzhou Huang, Ran He 0001
Pattern Recognit.5
2022 Cross-Spectral Iris Recognition by Learning Device-Specific Band
abstract
Cross-spectral recognition is still an open challenge in iris recognition. In cross-spectral iris recognition, there exist distinct device-specific bands between near-infrared (NIR) and visible (VIS) images, resulting in the distribution gap between samples from different spectra and thus severe degradation in recognition performance. To tackle this problem, we propose a new cross-spectral iris recognition method to learn spectral-invariant features by estimating device-specific bands. In the proposed method,GaborTridentNetwork (GTN) first utilizes the Gabor function’s priors to perceive iris textures under different spectra, and then codes the device-specific band as the residual component to assist the generation of spectral-invariant features. By investigating the device-specific band, GTN effectively reduces the impact of device-specific bands on identity features. Besides, we make three efforts to further reduce the distribution gap. First,SpectralAdversarialNetwork (SAN) adopts a class-level adversarial strategy to align feature distributions. Second,Sample-Anchor (SA) loss upgrades triplet loss by pulling samples to their class center and pushing away from other class centers. Third, we develop a higher-order alignment loss to measures the distribution gap according to space bases and distribution shapes. Extensive experiments on five iris datasets demonstrate the efficacy of our proposed method for cross-spectral iris recognition.
Jianze Wei, Yunlong Wang 0003, Yi Li 0018, Ran He 0001, Zhenan Sun
IEEE Trans. Circuits Syst. Video Technol.4
2022 Memory-Modulated Transformer Network for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) aims at matching face images across different domains. It is challenging due to the severe domain discrepancies and overfitting caused by small training datasets. Some researchers apply a “recognition via generation” strategy and propose to solve the problem by translating images from a given domain into the visual domain. However, in many HFR tasks such as near-infrablack HFR, there is no paiblack data, which makes it an unsupervised generation. Pose variations, background differences, and many other factors present challenges. Moreover, the generated results lack diversity since many previous works regard this image translation as a “one-to-one” generation task. Considering the information deficiency in the input images, we propose to formulate this image translation process as a “one-to-many” generation problem. Specifically, we introduce reference images to guide the generation process. We propose a memory module to explore the prototypical style patterns of the reference domain. After self-supervised updating, the memory items are attentively aggregated to represent the style information. Moreover, to subtly fuse the contents of input images with the style of reference images, we propose a novel style transformer module. Specifically, we crop the encoded input and reference feature maps into patches, and use the style transformer to establish long-range dependencies between the input and reference patches. Thus, the style of every input patch is transferblack based on those of the most relevant reference patches. Extensive experiments on multiple datasets for various HFR tasks, including NIR-VIS, thermal-VIS, sketch-photo, and gray-RGB, are conducted. The robustness and effectiveness of the proposed MMTN are demonstrated both quantitatively and qualitatively.
Mandi Luo, Haoxue Wu, Huaibo Huang, Weizan He, Ran He 0001
IEEE Trans. Inf. Forensics Secur.5
2022 Everybody's Talkin': Let Me Talk as You Want
abstract
We present a method to edit a target portrait footage by taking a sequence of audio as input to synthesize a photo-realistic video. This method is unique because it is highly dynamic. It does not assume a person-specific rendering network yet capable of translating one source audio into one random chosen video output within a set of speech videos. Instead of learning a highly heterogeneous and nonlinear mapping from audio to the video directly, we first factorize each target video frame into orthogonal parameter spaces,i.e., expression, geometry, and pose, via monocular 3D face reconstruction. Next, a recurrent network is introduced to translate source audio into expression parameters that are primarily related to the audio content. The audio-translated expression parameters are then used to synthesize a photo-realistic human subject in each video frame, with the movement of the mouth regions precisely mapped to the source audio. The geometry and pose parameters of the target human portrait are retained, therefore preserving the context of the original video footage. Finally, we introduce a novel video rendering network and a dynamic programming method to construct a temporally coherent and photo-realistic video. Extensive experiments demonstrate the superiority of our method over existing approaches. Our method is end-to-end learnable and robust to voice variations in the source audio.
Linsen Song, Wayne Wu, Chen Qian 0006, Ran He 0001, Chen Change Loy
IEEE Trans. Inf. Forensics Secur.4
2022 Towards More Discriminative and Robust Iris Recognition by Learning Uncertain Factors
abstract
The uncontrollable acquisition process limits the performance of iris recognition. In the acquisition process, various inevitable factors, including eyes, devices, and environment, hinder the iris recognition system from learning a discriminative identity representation. This leads to severe performance degradation. In this paper, we explore uncertain acquisition factors and propose uncertainty embedding (UE) and uncertainty-guided curriculum learning (UGCL) to mitigate the influence of acquisition factors. UE represents an iris image using a probabilistic distribution rather than a deterministic point (binary template or feature vector) that is widely adopted in iris recognition methods. Specifically, UE learns identity and uncertainty features from the input image, and encodes them as two independent components of the distribution, mean and variance. Based on this representation, an input image can be regarded as an instantiated feature sampled from the UE, and we can also generate various virtual features through sampling. UGCL is constructed by imitating the progressive learning process of newborns. Particularly, it selects virtual features to train the model in an easy-to-hard order at different training stages according to their uncertainty. In addition, an instance-level enhancement method is developed by utilizing local and global statistics to mitigate the data uncertainty from image noise and acquisition conditions in the pixel-level space. The experimental results on six benchmark iris datasets verify the effectiveness and generalization ability of the proposed method on same-sensor and cross-sensor recognition.
Jianze Wei, Huaibo Huang, Yunlong Wang 0003, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2021 ReMix: Towards Image-to-Image Translation With Limited Data
abstract
Image-to-image (I2I) translation methods based on generative adversarial networks (GANs) typically suffer from overfitting when limited training data is available. In this work, we propose a data augmentation method (ReMix) to tackle this issue. We interpolate training samples at the feature level and propose a novel content loss based on the perceptual relations among samples. The generator learns to translate the in-between samples rather than memorizing the training set, and thereby forces the discriminator to generalize. The proposed approach effectively reduces the ambiguity of generation and renders content-preserving results. The ReMix method can be easily incorporated into existing GAN models with minor modifications. Experimental results on numerous tasks demonstrate that GAN models equipped with the ReMix method achieve significant improvements.
Jie Cao 0002, Luanxuan Hou, Ming-Hsuan Yang 0001, Ran He 0001, Zhenan Sun
CVPR4
2021 Information Bottleneck Disentanglement for Identity Swapping
abstract
Improving the performance of face forgery detectors often requires more identity-swapped images of higher-quality. One core objective of identity swapping is to generate identity-discriminative faces that are distinct from the target while identical to the source. To this end, properly disentangling identity and identity-irrelevant information is critical and remains a challenging endeavor. In this work, we propose a novel information disentangling and swapping network, called InfoSwap, to extract the most expressive information for identity representation from a pre-trained face recognition model. The key insight of our method is to formulate the learning of disentangled representations as optimizing an information bottleneck tradeoff, in terms of finding an optimal compression of the pretrained latent features. Moreover, a novel identity contrastive loss is proposed for further disentanglement by requiring a proper distance between the generated identity and the target. While the most prior works have focused on using various loss functions to implicitly guide the learning of representations, we demonstrate that our model can provide explicit supervision for learning disentangled representations, achieving impressive performance in generating more identity-discriminative swapped faces.
Gege Gao, Huaibo Huang, Chaoyou Fu, Ran He 0001
CVPR5
2021 Memory Oriented Transfer Learning for Semi-Supervised Image Deraining
abstract
Deep learning based methods have shown dramatic improvements in image rain removal by using large-scale paired data of synthetic datasets. However, due to the various appearances of real rain streaks that may differ from those in the synthetic training data, it is challenging to directly extend existing methods to the real-world scenes. To address this issue, we propose a memory-oriented semi-supervised (MOSS) method which enables the network to explore and exploit the properties of rain streaks from both synthetic and real data. The key aspect of our method is designing an encoder-decoder neural network that is augmented with a self-supervised memory module, where items in the memory record the prototypical patterns of rain degradations and are updated in a self-supervised way. Consequently, the rainy styles can be comprehensively de-rived from synthetic or real-world degraded images without the need for clean labels. Furthermore, we present a self-training mechanism that attempts to transfer deraining knowledge from supervised rain removal to unsupervised cases. An additional target network, which is updated with an exponential moving average of the online deraining network, is utilized to produce pseudo-labels for unlabeled rainy images. Meanwhile, the deraining network is optimized with supervised objectives on both synthetic paired data and pseudo-paired noisy data. Extensive experiments show that the proposed method achieves more appealing results not only on limited labeled data but also on unlabeled real-world images than recent state-of-the-art methods.
Huaibo Huang, Aijing Yu, Ran He 0001
CVPR3
2021 FaceInpainter: High Fidelity Face Adaptation to Heterogeneous Domains
abstract
In this work, we propose a novel two-stage framework named FaceInpainter to implement controllable Identity-Guided Face Inpainting (IGFI) under heterogeneous domains. Concretely, by explicitly disentangling foreground and background of the target face, the first stage focuses on adaptive face fitting to the fixed background via a Styled Face Inpainting Network (SFI-Net), with 3D priors and texture code of the target, as well as identity factor of the source face. It is challenging to deal with the inconsistency between the new identity of the source and the original background of the target, concerning the face shape and appearance on the fused boundary. The second stage consists of a Joint Refinement Network (JR-Net) to refine the swapped face. It leverages AdaIN considering identity and multi-scale texture codes, for feature transformation of the decoded face from SFI-Net with facial occlusions. We adopt the contextual loss to implicitly preserve the attributes, encouraging face deformation and fewer texture distortions. Experimental results demonstrate that our approach handles high-quality identity adaptation to heterogeneous domains, exhibiting the competitive performance compared with state-of-the-art methods concerning both attribute and identity fidelity.
Jia Li 0044, Jie Cao 0002, Xingguang Song, Ran He 0001
CVPR5
2021 Pareidolia Face Reenactment
abstract
We present a new application direction named Pareidolia Face Reenactment, which is defined as animating a static illusory face to move in tandem with a human face in the video. For the large differences between pareidolia face reenactment and traditional human face reenactment, two main challenges are introduced, i.e., shape variance and texture variance. In this work, we propose a novel Parametric Unsupervised Reenactment Algorithm to tackle these two challenges. Specifically, we propose to decompose the reenactment into three catenate processes: shape modeling, motion transfer and texture synthesis. With the decomposition, we introduce three crucial components, i.e., Parametric Shape Modeling, Expansionary Motion Transfer and Unsupervised Texture Synthesizer, to overcome the problems brought by the remarkably variances on pareidolia faces. Extensive experiments show the superior performance of our method both qualitatively and quantitatively. Code, model and data are available on our project page1.
Linsen Song, Wayne Wu, Chaoyou Fu, Chen Qian 0006, Chen Change Loy, Ran He 0001
CVPR6
2021 Self-Augmented Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) is quite challenging due to the large discrepancy introduced by cross-domain face images. The limited number of paired face images results in a severe overfitting problem in existing methods. To tackle this issue, we proposes a novel self-augmentation method named Mixed Adversarial Examples and Logits Replay (MAELR). Concretely, we first generate adversarial examples, and mix them with clean examples in an interpolating way for data augmentation. Simultaneously, we extend the definition of the adversarial examples according to cross-domain problems. Benefiting from this extension, we can reduce domain discrepancy to extract domain-invariant features. We further propose a diversity preserving loss via logits replay, which effectively uses the discriminative features obtained on the large-scale VIS dataset. In this way, we improve the feature diversity that can not be obtained from mixed adversarial examples methods. Extensive experiments demonstrate that our method alleviates the over-fitting problem, thus significantly improving the recognition performance of HFR.
Zongcai Sun, Chaoyou Fu, Mandi Luo, Ran He 0001
IJCB4
2021 Contrastive Uncertainty Learning for Iris Recognition with Insufficient Labeled Samples
abstract
Cross-database recognition is still an unavoidable challenge when deploying an iris recognition system to a new environment. In the paper, we present a compromise problem that resembles the real-world scenario, named iris recognition with insufficient labeled samples. This new problem aims to improve the recognition performance by utilizing partially-or un-labeled data. To address the problem, we propose Contrastive Uncertainty Learning (CUL) by integrating the merits of uncertainty learning and contrastive self-supervised learning. CUL makes two efforts to learn a discriminative and robust feature representation. On the one hand, CUL explores the uncertain acquisition factors and adopts a probabilistic embedding to represent the iris image. In the probabilistic representation, the identity information and acquisition factors are disentangled into the mean and variance, avoiding the impact of uncertain acquisition factors on the identity information. On the other hand, CUL utilizes probabilistic embeddings to generate virtual positive and negative pairs. Then CUL builds its contrastive loss to group the similar samples closely and push the dissimilar samples apart. The experimental results demonstrate the effectiveness of the proposed CUL for iris recognition with insufficient labeled samples.
Jianze Wei, Ran He 0001, Zhenan Sun
IJCB2
2021 Visual-Semantic Transformer for Face Forgery Detection
abstract
This paper proposes a novel Visual-Semantic Transformer (VST) to detect face forgery based on semantic aware feature relations. In face images, intrinsic feature relations exist between different semantic parsing regions. We find that face forgery algorithms always change such relations. Therefore, we start the approach by extracting Contextual Feature Sequence (CFS) using a transformer encoder to make the best abnormal feature relation patterns. Meanwhile, images are segmented as soft face regions by a face parsing module. Then we merge the CFS and the soft face regions as Visual Semantic Sequences (VSS) representing features of semantic regions. The VSS is fed into the transformer decoder, in which the relations in the semantic region level are modeled. Our method achieved 99.58% accuracy on FF++(Raw) and 96.16% accuracy on Celeb-DF. Extensive experiments demonstrate that our framework outperforms or is comparable with state-of-the-art detection methods, especially towards unseen forgery methods.
Gengyun Jia, Huaibo Huang, Junxian Duan, Ran He 0001
IJCB5
2021 CM-NAS: Cross-Modality Neural Architecture Search for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared person re-identification (VI-ReID) aims to match cross-modality pedestrian images, breaking through the limitation of single-modality person ReID in dark environment. In order to mitigate the impact of large modality discrepancy, existing works manually design various two-stream architectures to separately learn modality-specific and modality-sharable representations. Such a manual design routine, however, highly depends on massive experiments and empirical practice, which is time consuming and labor intensive. In this paper, we systematically study the manually designed architectures, and identify that appropriately separating Batch Normalization (BN) layers is the key to bring a great boost towards cross-modality matching. Based on this observation, the essential objective is to find the optimal separation scheme for each BN layer. To this end, we propose a novel method, named Cross-Modality Neural Architecture Search (CM-NAS). It consists of a BN-oriented search space in which the standard optimization can be fulfilled subject to the cross-modality task. Equipped with the searched architecture, our method outperforms state-of-the-art counter-parts in both two benchmarks, improving the Rank-1/mAP by 6.70%/6.13% on SYSU-MM01 and by 12.17%/11.23% on RegDB. Code is released at https://github.com/JDAI-CV/CM-NAS.
Chaoyou Fu, Yibo Hu 0001, Xiang Wu 0001, Hailin Shi, Tao Mei 0001, Ran He 0001
ICCV6
2021 Invisible Backdoor Attack with Sample-Specific Triggers
abstract
Recently, backdoor attacks pose a new security threat to the training process of deep neural networks (DNNs). Attackers intend to inject hidden backdoors into DNNs, such that the attacked model performs well on benign samples, whereas its prediction will be maliciously changed if hidden backdoors are activated by the attacker-defined trigger. Existing backdoor attacks usually adopt the setting that triggers are sample-agnostic, i.e., different poisoned samples contain the same trigger, resulting in that the attacks could be easily mitigated by current backdoor defenses. In this work, we explore a novel attack paradigm, where backdoor triggers are sample-specific. In our attack, we only need to modify certain training samples with invisible perturbation, while not need to manipulate other training components (e.g., training loss, and model structure) as required in many existing attacks. Specifically, inspired by the recent advance in DNN-based image steganography, we generate sample-specific invisible additive noises as backdoor triggers by encoding an attacker-specified string into benign images through an encoder-decoder network. The mapping from the string to the target label will be generated when DNNs are trained on the poisoned dataset. Extensive experiments on benchmark datasets verify the effectiveness of our method in attacking models with or without defenses. The code will be available at https://github.com/yuezunli/ISSBA.
Yuezun Li, Yiming Li 0004, Baoyuan Wu, Longkang Li, Ran He 0001, Siwei Lyu
ICCV5
2021 Graph Information Bottleneck for Subgraph Recognition
Junchi Yu, Tingyang Xu, Yu Rong 0001, Yatao Bian, Junzhou Huang, Ran He 0001
ICLR6
2021 Selective Wavelet Attention Learning for Single Image Deraining
Huaibo Huang, Aijing Yu, Zhenhua Chai, Ran He 0001, Tieniu Tan
Int. J. Comput. Vis.4
2021 AutoDet: Pyramid Network Architecture Search for Object Detection
Zhihang Li, Teng Xi, Jingtuo Liu, Ran He 0001
Int. J. Comput. Vis.5
2021 LAMP-HQ: A Large-Scale Multi-pose High-Quality Database and Benchmark for NIR-VIS Face Recognition
Aijing Yu, Haoxue Wu, Huaibo Huang, Zhen Lei 0001, Ran He 0001
Int. J. Comput. Vis.5
2021 Coupled adversarial learning for semi-supervised heterogeneous face recognition
Ran He 0001, Yi Li 0018, Xiang Wu 0001, Lingxiao Song, Zhenhua Chai, Xiaolin Wei
Pattern Recognit.1
2021 High-Fidelity Face Manipulation With Extreme Poses and Expressions
abstract
Face manipulation has shown remarkable advances with the flourish of Generative Adversarial Networks. However, due to the difficulties of controlling structures and textures, it is challenging to model poses and expressions simultaneously, especially for the extreme manipulation at high-resolution. In this article, we propose a novel framework that simplifies face manipulation into two correlated stages: a boundary prediction stage and a disentangled face synthesis stage. The first stage models poses and expressions jointly via boundary images. Specifically, a conditional encoder-decoder network is employed to predict the boundary image of the target face in a semi-supervised way. Pose and expression estimators are introduced to improve the prediction performance. In the second stage, the predicted boundary image and the input face image are encoded into the structure and the texture latent space by two encoder networks, respectively. A proxy network and a feature threshold loss are further imposed to disentangle the latent space. Furthermore, due to the lack of high-resolution face manipulation databases to verify the effectiveness of our method, we collect a new high-quality Multi-View Face (MVF-HQ) database. It contains 120,283 images at 6000 × 4000 resolution from 479 identities with diverse poses, expressions, and illuminations. MVF-HQ is much larger in scale and much higher in resolution than publicly available high-resolution face manipulation databases. We will release MVF-HQ soon to push forward the advance of face manipulation. Qualitative and quantitative experiments on four databases show that our method dramatically improves the synthesis quality.
Chaoyou Fu, Yibo Hu 0001, Xiang Wu 0001, Guoli Wang 0004, Qian Zhang 0009, Ran He 0001
IEEE Trans. Inf. Forensics Secur.6
2021 FA-GAN: Face Augmentation GAN for Deformation-Invariant Face Recognition
abstract
Substantial improvements have been achieved in the field of face recognition due to the successful application of deep neural networks. However, existing methods are sensitive to both the quality and quantity of the training data. Despite the availability of large-scale datasets, the long tail data distribution induces strong biases in model learning. In this paper, we present a Face Augmentation Generative Adversarial Network (FA-GAN) to reduce the influence of imbalanced deformation attribute distributions. We propose to decouple these attributes from the identity representation with a novel hierarchical disentanglement module. Moreover, Graph Convolutional Networks (GCNs) are applied to recover geometric information by exploring the interrelations among local regions to guarantee the preservation of identities in face data augmentation. Extensive experiments on face reconstruction, face manipulation, and face recognition demonstrate the effectiveness and generalization ability of the proposed method.
Mandi Luo, Jie Cao 0002, Xin Ma 0031, Xiaoyu Zhang 0002, Ran He 0001
IEEE Trans. Inf. Forensics Secur.5
2021 Partial NIR-VIS Heterogeneous Face Recognition With Automatic Saliency Search
abstract
Near-infrared-visual (NIR-VIS) heterogeneous face recognition (HFR) aims to match NIR face images with the corresponding VIS ones. It is a challenging task due to the sensing gaps among different modalities. Occlusions in the input face images make the task extremely complex. To tackle these problems, we present a Saliency Search Network (SSN) to extract domain-invariant identity features. We propose to automatically search the efficient parts of face images in a modality-aware manner, and remove redundant information. Moreover, the searching process is guided by an information bottleneck network, which mitigates the overfitting problems caused by small datasets. Extensive experiments on both complete and partial NIR-VIS HFR on multiple datasets demonstrate the effectiveness and robustness of the proposed method to modality discrepancy and occlusions.
Mandi Luo, Xin Ma 0031, Zhihang Li, Jie Cao 0002, Ran He 0001
IEEE Trans. Inf. Forensics Secur.5
2020 Cross-Spectral Face Hallucination via Disentangling Independent Factors
abstract
The cross-sensor gap is one of the challenges that have aroused much research interests in Heterogeneous Face Recognition (HFR). Although recent methods have attempted to fill the gap with deep generative networks, most of them suffer from the inevitable misalignment between different face modalities. Instead of imaging sensors, the misalignment primarily results from facial geometric variations that are independent of the spectrum. Rather than building a monolithic but complex structure, this paper proposes a Pose Aligned Cross-spectral Hallucination (PACH) approach to disentangle the independent factors and deal with them in individual stages. In the first stage, an Unsupervised Face Alignment (UFA) module is designed to align the facial shapes of the near-infrared (NIR) images with those of the visible (VIS) images in a generative way, where UV maps are effectively utilized as the shape guidance. Thus the task of the second stage becomes spectrum translation with aligned paired data. We develop a Texture Prior Synthesis (TPS) module to achieve complexion control and consequently generate more realistic VIS images than existing methods. Experiments on three challenging NIR-VIS datasets verify the effectiveness of our approach in producing visually appealing images and achieving state-of-the-art performance in HFR.
Boyan Duan, Chaoyou Fu, Yi Li 0018, Xingguang Song, Ran He 0001
CVPR5
2020 PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer
abstract
In this paper, we address the makeup transfer task, which aims to transfer the makeup from a reference image to a source image. Existing methods have achieved promising progress in constrained scenarios, but transferring between images with large pose and expression differences is still challenging. Besides, they cannot realize customizable transfer that allows a controllable shade of makeup or specifies the part to transfer, which limits their applications. To address these issues, we propose Pose and expression robust Spatial-aware GAN (PSGAN). It first utilizes Makeup Distill Network to disentangle the makeup of the reference image as two spatial-aware makeup matrices. Then, Attentive Makeup Morphing module is introduced to specify how the makeup of a pixel in the source image is morphed from the reference image. With the makeup matrices and the source image, Makeup Apply Network is used to perform makeup transfer. Our PSGAN not only achieves state-of-the-art results even when large pose and expression differences exist but also is able to perform partial and shade-controllable makeup transfer. Both the code and a newly collected dataset containing facial images with various poses and expressions will be available at https://github.com/wtjiang98/PSGAN.
Si Liu 0001, Chen Gao 0005, Jie Cao 0002, Ran He 0001, Jiashi Feng, Shuicheng Yan
CVPR5
2020 GP-NAS: Gaussian Process Based Neural Architecture Search
abstract
Neural architecture search (NAS) advances beyond the state-of-the-art in various computer vision tasks by automating the designs of deep neural networks. In this paper, we aim to address three important questions in NAS: (1) How to measure the correlation between architectures and their performances? (2) How to evaluate the correlation between different architectures? (3) How to learn these correlations with a small number of samples? To this end, we first model these correlations from a Bayesian perspective. Specifically, by introducing a novel Gaussian Process based NAS (GP-NAS) method, the correlations are modeled by the kernel function and mean function. The kernel function is also learnable to enable adaptive modeling for complex correlations in different search spaces. Furthermore, by incorporating a mutual information based sampling method, we can theoretically ensure the high-performance architecture with only a small set of samples. After addressing these problems, training GP-NAS once enables direct performance prediction of any architecture in different scenarios and may obtain efficient networks for different deployment platforms. Extensive experiments on both image classification and face recognition tasks verify the effectiveness of our algorithm.
Zhihang Li, Teng Xi, Jiankang Deng, Shengzhao Wen, Ran He 0001
CVPR6
2020 Informative Sample Mining Network for Multi-domain Image-to-Image Translation
Jie Cao 0002, Huaibo Huang, Yi Li 0018, Ran He 0001, Zhenan Sun
ECCV (19)4
2020 TF-NAS: Rethinking Three Search Freedoms of Latency-Constrained Differentiable Neural Architecture Search
Yibo Hu 0001, Xiang Wu 0001, Ran He 0001
ECCV (15)3
2020 Hierarchical Face Aging Through Disentangled Latent Characteristics
Peipei Li 0002, Huaibo Huang, Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun
ECCV (3)5
2020 A Balanced and Uncertainty-Aware Approach for Partial Domain Adaptation
Jian Liang 0001, Yunbo Wang, Dapeng Hu, Ran He 0001, Jiashi Feng
ECCV (11)4
2020 MEAD: A Large-Scale Audio-Visual Dataset for Emotional Talking-Face Generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian 0006, Ran He 0001, Yu Qiao 0001, Chen Change Loy
ECCV (21)7
2020 Unsupervised Contrastive Photo-to-Caricature Translation based on Auto-distortion
abstract
Photo-to-caricature translation aims to synthesize the caricature as a rendered image exaggerating the features through sketching, pencil strokes, or other artistic drawings. Style rendering and geometry deformation are the most important aspects in photo-to-caricature translation task. To take both into consideration, we propose an unsupervised contrastive photo-to-caricature translation architecture. Considering the intuitive artifacts in the existing methods, we propose a contrastive style loss for style rendering to enforce the similarity between the style of rendered photo and the caricature, and simultaneously enhance its discrepancy to the photos. To obtain an exaggerating deformation in an unpaired/unsupervised fashion, we propose a Distortion Prediction Module (DPM) to predict a set of displacements vectors for each input image while fixing some controlling points, followed by the thin plate spline interpolation for warping. The model is trained on unpaired photo and caricature while can offer bidirectional synthesizing via inputting either a photo or a caricature. Extensive experiments demonstrate that the proposed model is effective to generate hand-drawn like caricatures compared with existing competitors.
Yuhe Ding, Xin Ma 0031, Mandi Luo, Aihua Zheng, Ran He 0001
ICPR5
2020 $P^{2}$ Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
abstract
The target of human pose estimation is to determine the body parts and joint locations of persons in the image. Angular changes, motion blur and occlusion in the natural scenes make this task challenging, while some joints are more difficult to be detected than others. In this paper, we propose an augmented Parallel-Pyramid Net ( P2Net) with feature refinement by dilated bottleneck and attention module. During data preprocessing, we proposed a differentiable auto data augmentation ( DA2) method. We formulate the problem of searching data augmentaion policy in a differentiable form, so that the optimal policy setting can be easily updated by back propagation during training. DA2improves the training efficiency. A parallel-pyramid structure is followed to compensate the information loss introduced by the network. We innovate two fusion structures, i.e. Parallel Fusion and Progressive Fusion, to process pyramid features from backbone network. Both fusion structures leverage the advantages of spatial information affluence at high resolution and semantic comprehension at low resolution effectively. We propose a refinement stage for the pyramid features to further boost the accuracy of our network. By introducing dilated bottleneck and attention module, we increase the receptive field for the features with limited complexity and tune the importance to different feature channels. To further refine the feature maps after completion of feature extraction stage, an Attention Module ( AM) is defined to extract weighted features from different scale feature maps generated by the parallel-pyramid structure. Compared with the traditional up-sampling refining, AM can better capture the relationship between channels. Experiments corroborate the effectiveness of our proposed method. Notably, our method achieves the best performance on the challenging MSCOCO and MPII datasets.
Luanxuan Hou, Jie Cao 0002, Haifeng Shen, Jian Tang 0008, Ran He 0001
ICPR6
2020 Free-Form Image Inpainting via Contrastive Attention Network
abstract
Most deep learning based image inpainting approaches adopt autoencoder or its variants to fill missing regions in images. Encoders are usually utilized to learn powerful representational spaces, which are important for dealing with sophisticated learning tasks. Specifically, in image inpainting tasks, masks with any shapes can appear anywhere in images (i.e., free-form masks) which form complex patterns. It is difficult for encoders to capture such powerful representations under this complex situation. To tackle this problem, we propose a self-supervised Siamese inference network to improve the robustness and generalization. It can encode contextual semantics from full resolution images and obtain more discriminative representations. we further propose a multi-scale decoder with a novel dual attention fusion module (DAF), which can combine both the restored and known regions in a smooth way. This multi-scale architecture is benefit for decoding discriminative representations learned by encoders into images layer by layer. In this way, unknown regions will be filled naturally from outside to inside. Qualitative and quantitative experiments on multiple datasets, including facial and natural datasets (i.e., Celeb-HQ, Pairs Street View, Places2 and ImageNet), demonstrate that our proposed method outperforms state-of-the-art methods in generating high-quality inpainting results.
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Zhenhua Chai, Xiaolin Wei, Ran He 0001
ICPR6
2020 Attentional Wavelet Network for Traditional Chinese Painting Transfer
abstract
Traditional Chinese paintings pay more attention to `Gongbi' and `Xieyi' in artworks, which raises a challenging task to generate Chinese paintings from photos. `Xieyi' creates high-level conception for paintings, while `Gongbi' refers to portraying local details in paintings. This paper proposes an attentional wavelet network for photo to Chinese painting transferring. We first introduce wavelets to obtain high-level conception and local details in Chinese paintings via 2-D haar wavelet transform. Moreover, we design high-level transform stream and local enhancement stream to dispose high frequencies and low frequency respectively. Furthermore, we exploit self-attention mechanism to compatibly pick up high-level information which is used to remedy the missing details when reconstructing the Chinese painting. To advance our experiment, we set up a new dataset named P2ADataset, with diverse photos and Chinese paintings on famous mountains around China. Experimental results comparing with the state-of-the-art style transferring algorithms verify the effectiveness of the proposed method. We will release the codes and data to the public.
Rui Wang 0124, Huaibo Huang, Aihua Zheng, Ran He 0001
ICPR4
2020 Exemplar Guided Cross-Spectral Face Hallucination via Mutual Information Disentanglement
abstract
Recently, many Near infrared-visible (NIR-VIS) heterogeneous face recognition (HFR) methods have been proposed in the community. But it remains a challenging problem because of the sensing gap along with large pose variations. In this paper, we propose an Exemplar Guided Cross-Spectral Face Hallucination (EGCH) to reduce the domain discrepancy through disentangled representation learning. For each modality, EGCH contains a spectral encoder as well as a structure encoder to disentangle spectral and structure representation, respectively. It also contains a traditional generator that reconstructs the input from the above two representations, and a structure generator that predicts the facial parsing map from the structure representation. Besides, mutual information minimization and maximization are conducted to boost disentanglement and make representations adequately expressed. Then the translation is built on structure representations between two modalities. Provided with the transformed NIR structure representation and original VIS spectral representation, EGCH is capable to produce high-fidelity VIS images that preserve the topology structure of the input NIR while transfer the spectral information of an arbitrary VIS exemplar. Extensive experiments demonstrate that the proposed method achieves more promising results both qualitatively and quantitatively than the state-of-the-art NIR-VIS methods.
Haoxue Wu, Huaibo Huang, Aijing Yu, Jie Cao 0002, Zhen Lei 0001, Ran He 0001
ICPR6
2020 Talking Face Generation via Learning Semantic and Temporal Synchronous Landmarks
abstract
Given a speech clip and facial image, the goal of talking face generation is to synthesize a talking face video with accurate mouth synchronization and natural face motion. Recent progress has proven the effectiveness of the landmarks as the intermediate information during talking face generation. However, the large gap between audio and visual modalities makes the prediction of landmarks challenging and limits generation ability. This paper proposes a semantic and temporal synchronous landmark learning method for talking face generation. First, we propose to introduce a word detector to enforce richer semantic information. Then, we propose to preserve the temporal synchronization and consistency between landmarks and audio via the proposed temporal residual loss. Lastly, we employ a U-Net generation network with adaptive reconstruction loss to generate facial images for the predicted landmarks. Experimental results on two benchmark datasets LRW and GRID demonstrate the effectiveness of our model compared to the state-of-the-art methods of talking face generation.
Aihua Zheng, Feixia Zhu, Mandi Luo, Ran He 0001
ICPR5
2020 Image Inpainting with Contrastive Relation Network
abstract
Image inpainting faces the challenging issue of the requirements on structure reasonableness and texture coherence. In this paper, we propose a two-stage inpainting framework to address this issue. The basic idea is to address the two requirements in two separate stages. Completed segmentation of the corrupted image is firstly predicted through segmentation reconstruction network, while fine-grained image details are restored in the second stage through an image generator. The two stages are connected in series as the image details are generated under the guidance of completed segmentation map that predicted in the first stage. Specifically, in the second stage, we propose a novel graph-based relation network to model the relationship existed in corrupted image. In relation network, both intra-relationship for pixels in the same semantic region and inter-relationship between different semantic parts are considered, improving the consistency and compatibility of image textures. Besides, contrastive loss is designed to facilitate the relation network training. Such a framework not only simplifies the inpainting problem directly, but also exploits the relationship in corrupted image explicitly. Extensive experiments on various public datasets quantitatively and qualitatively demonstrate the superiority of our approach compared with the state-of-the-art.
Xiaoqiang Zhou, Junjie Li 0002, Zilei Wang, Ran He 0001, Tieniu Tan
ICPR4
2020 Let's Play Music: Audio-Driven Performance Video Generation
abstract
We propose a new task named Audio-driven Performance Video Generation (APVG), which aims to synthesize the video of a person playing a certain instrument guided by a given music audio clip. It is a challenging task to generate the high-dimensional temporal consistent videos from low-dimensional audio modality. In this paper, we propose a multi-staged framework to generate realistic and synchronized performance video from given music. Firstly, we provide both global appearance and local spatial information by generating the coarse videos and keypoints of body and hands from a given music respectively. Then, we propose to transform the generated keypoints to heatmap via a differentiable space transformer, since the heatmap provides more spatial information but is harder to generate directly from audio. Finally, we propose a Structured Temporal UNet (STU) to extract both intra-frame structured information and interframe temporal consistency. They are obtained via graph-based structure module, and CNN-GRU based high-level temporal module respectively for final video generation. Comprehensive experiments validate the effectiveness of our proposed framework.
Yi Li 0018, Feixia Zhu, Aihua Zheng, Ran He 0001
ICPR5
2020 Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence Learning
abstract
Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on either disentangling the information in a single image or learning temporal information between frames. However, cross-modality coherence between audio and video information has not been well addressed during synthesis. In this paper, we propose a novel arbitrary talking face generation framework by discovering the audio-visual coherence via the proposed Asymmetric Mutual Information Estimator (AMIE). In addition, we propose a Dynamic Attention (DA) block by selectively focusing the lip area of the input image during the training stage, to further enhance lip synchronization. Experimental results on benchmark LRW dataset and GRID dataset transcend the state-of-the-art methods on prevalent metrics with robust high-resolution synthesizing on gender and pose variations.
Huaibo Huang, Yi Li 0018, Aihua Zheng, Ran He 0001
IJCAI5
2020 InteractGAN: Learning to Generate Human-Object Interaction
abstract
Compared with the widely studied Human-Object Interaction DE-Tection (HOI-DET), no effort has been devoted to its inverse problem, i.e. to generate an HOI scene image according to the given relationship triplet , to our best knowledge. We term this new task "Human-Object Interaction Image Generation" (HOI-IG). HOI-IG is a research-worthy task with great application prospects, such as online shopping, film production and interactive entertainment. In this work, we introduce an Interact-GAN to solve this challenging task. Our method is composed of two stages: (1) manipulating the posture of a given human image conditioned on a predicate. (2) merging the transformed human image and object image to one realistic scene image while satisfying the ir expected relative position and ratio. Besides, to address the large spatial misalignment issue caused by fusing two images content with reasonable spatial layout, we propose a Relation-based Spatial Transformer Network (RSTN) to adaptively process the images conditioned on their interaction. Extensive experiments on two challenging datasets demonstrate the effectiveness and superiority of our approach. We advocate for the image generation community to draw more attention to the new Human-Object Interaction Image Generation problem. To facilitate future research, our project will be released at: http://colalab.org/projects/InteractGAN.
Chen Gao 0005, Si Liu 0001, Defa Zhu, Jie Cao 0002, Haoqian He, Ran He 0001, Shuicheng Yan
ACM Multimedia7
2020 Beautify As You Like
abstract
Customizable makeup transfer, which aims to transfer the makeup from an arbitrary reference face to a source face, is widely demanded in many applications such as short video platforms and online meeting applications. However, existing methods are neither user-friendly nor sufficiently fast. In this demo, we present the first fast makeup transfer system named as Fast Pose and expression robust Spatial-Aware GAN (FPSGAN). With a novel Attentive Makeup Morphing (AMM) module, FPSGAN is robust to face pose and expression. Moreover, it can achieve shade-controllable and partial makeup, improving the system's user-friendliness. In addition, FPSGAN is light-weighted and fast. To sum up, FPSGAN is the first fast customizable makeup transfer system to enable users to beautify themselves as they like.
Si Liu 0001, Chen Gao 0005, Ran He 0001, Bo Li 0006, Shuicheng Yan
ACM Multimedia4
2020 Dual-Structure Disentangling Variational Generation for Data-Limited Face Parsing
abstract
Deep learning based face parsing methods have attained state-of-the-art performance in recent years. Their superior performance heavily depends on the large-scale annotated training data. However, it is expensive and time-consuming to construct a large-scale pixel-level manually annotated dataset for face parsing. To alleviate this issue, we propose a novel Dual-Structure Disentangling Variational Generation (D2VG) network. Benefiting from the interpretable factorized latent disentanglement in VAE, D2VG can learn a joint structural distribution of facial image and its corresponding parsing map. Owing to these, it can synthesize large-scale paired face images and parsing maps from a standard Gaussian distribution. Then, we adopt both manually annotated and synthesized data to train a face parsing model in a supervised way. Since there are inaccurate pixel-level labels in synthesized parsing maps, we introduce a coarseness-tolerant learning algorithm, to effectively handle these noisy or uncertain labels. In this way, we can significantly boost the performance of face parsing. Extensive quantitative and qualitative results on HELEN, CelebAMask-HQ and LaPa demonstrate the superiority of our methods.
Peipei Li 0002, Yinglu Liu, Hailin Shi, Xiang Wu 0001, Yibo Hu 0001, Ran He 0001, Zhenan Sun
ACM Multimedia6
2020 AOT: Appearance Optimal Transport Based Identity Swapping for Forgery Detection
abstract
Recent studies have shown that the performance of forgery detection can be improved with diverse and challenging Deepfakes datasets. However, due to the lack of Deepfakes datasets with large variance in appearance, which can be hardly produced by recent identity swapping methods, the detection algorithm may fail in this situation. In this work, we provide a new identity swapping algorithm with large differences in appearance for face forgery detection. The appearance gaps mainly arise from the large discrepancies in illuminations and skin colors that widely exist in real-world scenarios. However, due to the difficulties of modeling the complex appearance mapping, it is challenging to transfer fine-grained appearances adaptively while preserving identity traits. This paper formulates appearance mapping as an optimal transport problem and proposes an Appearance Optimal Transport model (AOT) to formulate it in both latent and pixel space. Specifically, a relighting generator is designed to simulate the optimal transport plan. It is solved via minimizing the Wasserstein distance of the learned features in the latent space, enabling better performance and less computation than conventional optimization. To further refine the solution of the optimal transport plan, we develop a segmentation game to minimize the Wasserstein distance in the pixel space. A discriminator is introduced to distinguish the fake parts from a mix of real and fake image patches. Extensive experiments reveal that the superiority of our method when compared with state-of-the-art methods and the ability of our generated data to improve the performance of face forgery detection.
Chaoyou Fu, Qianyi Wu, Wayne Wu, Chen Qian 0006, Ran He 0001
NeurIPS6
2020 Towards High Fidelity Face Frontalization in the Wild
Jie Cao 0002, Yibo Hu 0001, Hongwen Zhang 0001, Ran He 0001, Zhenan Sun
Int. J. Comput. Vis.4
2020 Disentangled Representation Learning of Makeup Portraits in the Wild
Yi Li 0018, Huaibo Huang, Jie Cao 0002, Ran He 0001, Tieniu Tan
Int. J. Comput. Vis.4
2020 A General Framework for Deep Supervised Discrete Hashing
Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan
Int. J. Comput. Vis.3
2020 Learning an Evolutionary Embedding via Massive Knowledge Distillation
Xiang Wu 0001, Ran He 0001, Yibo Hu 0001, Zhenan Sun
Int. J. Comput. Vis.2
2020 A Survey of Deep Facial Attribute Analysis
Xin Zheng 0008, Yanqing Guo, Huaibo Huang, Yi Li 0018, Ran He 0001
Int. J. Comput. Vis.5
2020 Adversarial Cross-Spectral Face Completion for NIR-VIS Face Recognition
abstract
Near infrared-visible (NIR-VIS) heterogeneous face recognition refers to the process of matching NIR to VIS face images. Current heterogeneous methods try to extend VIS face recognition methods to the NIR spectrum by synthesizing VIS images from NIR images. However, due to the self-occlusion and sensing gap, NIR face images lose some visible lighting contents so that they are always incomplete compared to VIS face images. This paper models high-resolution heterogeneous face synthesis as a complementary combination of two components: a texture inpainting component and a pose correction component. The inpainting component synthesizes and inpaints VIS image textures from NIR image textures. The correction component maps any pose in NIR images to a frontal pose in VIS images, resulting in paired NIR and VIS textures. A warping procedure is developed to integrate the two components into an end-to-end deep network. A fine-grained discriminator and a wavelet-based discriminator are designed to improve visual quality. A novel 3D-based pose correction loss, two adversarial losses, and a pixel loss are imposed to ensure synthesis results. We demonstrate that by attaching the correction component, we can simplify heterogeneous face synthesis from one-to-many unpaired image translation to one-to-one paired image translation, and minimize the spectral and pose discrepancy during heterogeneous recognition. Extensive experimental results show that our network not only generates high-resolution VIS face images but also facilitates the accuracy improvement of heterogeneous face recognition.
Ran He 0001, Jie Cao 0002, Lingxiao Song, Zhenan Sun, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Learning disentangling and fusing networks for face completion under structured occlusions
Zhihang Li, Yibo Hu 0001, Ran He 0001, Zhenan Sun
Pattern Recognit.3
2020 Deep label refinement for age estimation
Peipei Li 0002, Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun
Pattern Recognit.4
2020 BLAN: Bi-directional ladder attentive network for facial attribute prediction
Xin Zheng 0008, Huaibo Huang, Yanqing Guo, Bo Wang 0024, Ran He 0001
Pattern Recognit.5
2020 Recurrent Prediction With Spatio-Temporal Attention for Crowd Attribute Recognition
abstract
Crowd attribute recognition is a challenging task for crowd video understanding because a crowd video often contains multiple attributes from various types. Traditional deep learning-based methods directly treat this recognition problem as a multiple binary classification problem and represent the video by vectorizing and fusing the separately learned spatial and temporal features in the fully connected layers. Therefore, the correlations between these attributes may not be well captured. In this paper, a bidirectional recurrent prediction model with a semantic-aware attention mechanism is proposed to explore the spatio-temporal and semantic relations between the attributes for more accurate recognition. The ConvLSTM is introduced for feature representation to capture the spatio-temporal structure of the crowd videos and facilitate the visual attention. The bidirectional recurrent attention module is proposed for sequential attribute prediction by associating each subcategory attributes to corresponding semantic-related regions iteratively. The experiments and evaluations on the challenging WWW crowd video dataset not only show that our approach significantly outperforms the state-of-the-art methods but also verify that our approach can effectively capture the spatio-temporal and semantic relations of the crowd attributes.
Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.3
2020 Progressively Refined Face Detection Through Semantics-Enriched Representation Learning
abstract
Feature pyramids aim to learn multi-scale representations for detecting faces over various scales. However, they often lack adequate context over different scales, especially when there are many tiny faces in the wild. In this paper, we propose an attention-guided semantically enriched feature aggregation framework to learn a feature pyramid with rich semantics at all scales for face detection. Specifically, high-level abstract features are directly integrated into low-level representations by skip connections to retain as much semantic as possible. In addition, an attention mechanism is employed as a gate to emphasize relevant features and suppress useless features during feature fusion. Inspired by human visual perception of tiny faces, we specially design a deep progressive refined loss (DPRL) to effectively facilitate feature learning. According to the above principles, we design and investigate various feature pyramid frameworks through extensive experiments. Finally, two typical structures named Centralized Attention Feature (CAF) and Distributed Attention Feature (DAF) are proposed for face detection, which are in-place and end-to-end trainable. Extensive experiments across different aggregation architectures on four challenging face detection benchmarks demonstrate the superiority of our framework over state-of-the-art methods.
Zhihang Li, Xu Tang 0007, Xiang Wu 0001, Jingtuo Liu, Ran He 0001
IEEE Trans. Inf. Forensics Secur.5
2019 Disentangled Variational Representation for Heterogeneous Face Recognition
abstract
Visible (VIS) to near infrared (NIR) face matching is a challenging problem due to the significant domain discrepancy between the domains and a lack of sufficient data for training cross-modal matching algorithms. Existing approaches attempt to tackle this problem by either synthesizing visible faces from NIR faces, extracting domain-invariant features from these modalities, or projecting heterogeneous data onto a common latent space for cross-modal matching. In this paper, we take a different approach in which we make use of the Disentangled Variational Representation (DVR) for crossmodal matching. First, we model a face representation with an intrinsic identity information and its within-person variations. By exploring the disentangled latent variable space, a variational lower bound is employed to optimize the approximate posterior for NIR and VIS representations. Second, aiming at obtaining more compact and discriminative disentangled latent space, we impose a minimization of the identity information for the same subject and a relaxed correlation alignment constraint between the NIR and VIS modality variations. An alternative optimization scheme is proposed for the disentangled variational representation part and the heterogeneous face recognition network part. The mutual promotion between these two parts effectively reduces the NIR and VIS domain discrepancy and alleviates over-fitting. Extensive experiments on three challenging NIR-VIS heterogeneous face recognition databases demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods.
Xiang Wu 0001, Huaibo Huang, Vishal M. Patel, Ran He 0001, Zhenan Sun
AAAI4
2019 Visual-Semantic Graph Reasoning for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition in surveillance is a challenging task due to poor image quality, significant appearance variations and diverse spatial distribution of different attributes. This paper treats pedestrian attribute recognition as a sequential attribute prediction problem and proposes a novel visual-semantic graph reasoning framework to address this problem. Our framework contains a spatial graph and a directed semantic graph. By performing reasoning using the Graph Convolutional Network (GCN), one graph captures spatial relations between regions and the other learns potential semantic relations between attributes. An end-to-end architecture is presented to perform mutual embedding between these two graphs to guide the relational learning for each other. We verify the proposed framework on three large scale pedestrian attribute datasets including PETA, RAP, and PA100k. Experiments show superiority of the proposed method over state-of-the-art methods and effectiveness of our joint GCN structures for sequential attribute prediction.
Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang
AAAI3
2019 Geometry-Aware Face Completion and Editing
abstract
Face completion is a challenging generation task because it requires generating visually pleasing new pixels that are semantically consistent with the unmasked face region. This paper proposes a geometry-aware Face Completion and Editing NETwork (FCENet) by systematically studying facial geometry from the unmasked region. Firstly, a facial geometry estimator is learned to estimate facial landmark heatmaps and parsing maps from the unmasked face image. Then, an encoder-decoder structure generator serves to complete a face image and disentangle its mask areas conditioned on both the masked face image and the estimated facial geometry images. Besides, since low-rank property exists in manually labeled masks, a low-rank regularization term is imposed on the disentangled masks, enforcing our completion network to manage occlusion area with various shape and size. Furthermore, our network can generate diverse results from the same masked input by modifying estimated facial geometry, which provides a flexible mean to edit the completed face appearance. Extensive experimental results qualitatively and quantitatively demonstrate that our network is able to generate visually pleasing face completion results and edit face attributes as well.
Linsen Song, Jie Cao 0002, Lingxiao Song, Yibo Hu 0001, Ran He 0001
AAAI5
2019 Distant Supervised Centroid Shift: A Simple and Efficient Approach to Visual Domain Adaptation
abstract
Conventional domain adaptation methods usually resort to deep neural networks or subspace learning to find invariant representations across domains. However, most deep learning methods highly rely on large-size source domains and are computationally expensive to train, while subspace learning methods always have a quadratic time complexity that suffers from the large domain size. This paper provides a simple and efficient solution, which could be regarded as a well-performing baseline for domain adaptation tasks. Our method is built upon the nearest centroid classifier, seeking a subspace where the centroids in the target domain are moderately shifted from those in the source domain. Specifically, we design a unified objective without accessing the source domain data and adopt an alternating minimization scheme to iteratively discover the pseudo target labels, invariant subspace, and target centroids. Besides its privacy-preserving property (distant supervision), the algorithm is provably convergent and has a promising linear time complexity. In addition, the proposed method can be readily extended to multi-source setting and domain generalization, and it remarkably enhances popular deep adaptation methods by borrowing the learned transferable features. Extensive experiments on several benchmarks including object, digit, and face recognition datasets validate that our methods yield state-of-the-art results in various domain adaptation tasks.
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan
CVPR2
2019 M2FPA: A Multi-Yaw Multi-Pitch High-Quality Dataset and Benchmark for Facial Pose Analysis
abstract
Facial images in surveillance or mobile scenarios often have large view-point variations in terms of pitch and yaw angles. These jointly occurred angle variations make face recognition challenging. Current public face databases mainly consider the case of yaw variations. In this paper, a new large-scale Multi-yaw Multi-pitch high-quality database is proposed for Facial Pose Analysis (M2FPA), including face frontalization, face rotation, facial pose estimation and pose-invariant face recognition. It contains 397,544 images of 229 subjects with yaw, pitch, attribute, illumination and accessory. M2FPA is the most comprehensive multi-view face database for facial pose analysis. Further, we provide an effective benchmark for face frontalization and pose-invariant face recognition on M2FPA with several state-of-the-art methods, including DR-GAN, TP-GAN and CAPG-GAN. We believe that the new database and benchmark can significantly push forward the advance of facial pose analysis in real-world applications. Moreover, a simple yet effective parsing guided discriminator is introduced to capture the local consistency during GAN optimization. Extensive quantitative and qualitative results on M2FPA and Multi-PIE demonstrate the superiority of our face frontalization method. Baseline results for both face synthesis and face recognition from state-of-the-art methods demonstrate the challenge offered by this new database.
Peipei Li 0002, Xiang Wu 0001, Yibo Hu 0001, Ran He 0001, Zhenan Sun
ICCV4
2019 Make a Face: Towards Arbitrary High Fidelity Face Manipulation
abstract
Recent studies have shown remarkable success in face manipulation task with the advance of GANs and VAEs paradigms, but the outputs are sometimes limited to low-resolution and lack of diversity. In this work, we propose Additive Focal Variational Auto-encoder (AF-VAE), a novel approach that can arbitrarily manipulate high-resolution face images using a simple yet effective model and only weak supervision of reconstruction and KL divergence losses. First, a novel additive Gaussian Mixture assumption is introduced with an unsupervised clustering mechanism in the structural latent space, which endows better disentanglement and boosts multi-modal representation with external memory. Second, to improve the perceptual quality of synthesized results, two simple strategies in architecture design are further tailored and discussed on the behavior of Human Visual System (HVS) for the first time, allowing for fine control over the model complexity and sample quality. Human opinion studies and new state-of-the-art Inception Score (IS) / Frechet Inception Distance (FID) demonstrate the superiority of our approach over existing algorithms, advancing both the fidelity and extremity of face manipulation task.
Shengju Qian, Kwan-Yee Lin, Wayne Wu, Yangxiaokang Liu, Fumin Shen, Chen Qian 0006, Ran He 0001
ICCV8
2019 Neurons Merging Layer: Towards Progressive Redundancy Reduction for Deep Supervised Hashing
abstract
Deep supervised hashing has become an active topic in information retrieval. It generates hashing bits by the output neurons of a deep hashing network. During binary discretization, there often exists much redundancy between hashing bits that degenerates retrieval performance in terms of both storage and accuracy. This paper proposes a simple yet effective Neurons Merging Layer (NMLayer) for deep supervised hashing. A graph is constructed to represent the redundancy relationship between hashing bits that is used to guide the learning of a hashing network. Specifically, it is dynamically learned by a novel mechanism defined in our active and frozen phases. According to the learned relationship, the NMLayer merges the redundant neurons together to balance the importance of each output neuron. Moreover, multiple NMLayers are progressively trained for a deep hashing network to learn a more compact hashing code from a long redundant code. Extensive experiments on four datasets demonstrate that our proposed method outperforms state-of-the-art hashing methods.
Chaoyou Fu, Liangchen Song, Xiang Wu 0001, Guoli Wang 0004, Ran He 0001
IJCAI5
2019 Pedestrian Attribute Recognition by Joint Visual-semantic Reasoning and Knowledge Distillation
abstract
Pedestrian attribute recognition in surveillance is a challenging task in computer vision due to significant pose variation, viewpoint change and poor image quality. To achieve effective recognition, this paper presents a graph-based global reasoning framework to jointly model potential visual-semantic relations of attributes and distill auxiliary human parsing knowledge to guide the relational learning. The reasoning framework models attribute groups on a graph and learns a projection function to adaptively assign local visual features to the nodes of the graph. After feature projection, graph convolution is utilized to perform global reasoning between the attribute groups to model their mutual dependencies. Then, the learned node features are projected back to visual space to facilitate knowledge transfer. An additional regularization term is proposed by distilling human parsing knowledge from a pre-trained teacher model to enhance feature representations. The proposed framework is verified on three large scale pedestrian attribute datasets including PETA, RAP, and PA-100k. Experiments show that our method achieves state-of-the-art results.
Qiaozhe Li, Xin Zhao 0012, Ran He 0001, Kaiqi Huang
IJCAI3
2019 Pose-preserving Cross Spectral Face Hallucination
abstract
To narrow the inherent sensing gap in heterogeneous face recognition (HFR), recent methods have resorted to generative models and explored the ?recognition via generation? framework. Even though, it remains a very challenging task to synthesize photo-realistic visible faces (VIS) from near-infrared (NIR) images especially when paired training data are unavailable. We present an approach to avert the data misalignment problem and faithfully preserve pose, expression and identity information during cross-spectral face hallucination. At the pixel level, we introduce an unsupervised attention mechanism to warping that is jointly learned with the generator to derive pixel-wise correspondence from unaligned data. At the image level, an auxiliary generator is employed to facilitate the learning of mapping from NIR to VIS domain. At the domain level, we first apply the mutual information constraint to explicitly measure the correlation between domains and thus benefit synthesis. Extensive experiments on three heterogeneous face datasets demonstrate that our approach not only outperforms current state-of-the-art HFR methods but also produce visually appealing results at a high resolution.
Junchi Yu, Jie Cao 0002, Yi Li 0018, Xiaofei Jia, Ran He 0001
IJCAI5
2019 Learning Disentangled Representation for Cross-Modal Retrieval with Deep Mutual Information Estimation
abstract
Cross-modal retrieval has become a hot research topic in recent years for its theoretical and practical significance. This paper proposes a new technique for learning such deep visual-semantic embedding that is more effective and interpretable for cross-modal retrieval. The proposed method employs a two-stage strategy to fulfill the task. In the first stage, deep mutual information estimation is incorporated into the objective to maximize the mutual information between the input data and its embedding. In the second stage, an expelling branch is added to the network to disentangle the modality-exclusive information from the learned representations. This helps to reduce the impact of modality-exclusive information to the common subspace representation as well as improve the interpretability of the learned feature. Extensive experiments on two large-scale benchmark datasets demonstrate that our method can learn better visual-semantic embedding and achieve state-of-the-art cross-modal retrieval results.
Weikuo Guo, Huaibo Huang, Xiangwei Kong 0001, Ran He 0001
ACM Multimedia4
2019 Dual Variational Generation for Low Shot Heterogeneous Face Recognition
abstract
Heterogeneous Face Recognition (HFR) is a challenging issue because of the large domain discrepancy and a lack of heterogeneous data. This paper considers HFR as a dual generation problem, and proposes a novel Dual Variational Generation (DVG) framework. It generates large-scale new paired heterogeneous images with the same identity from noise, for the sake of reducing the domain gap of HFR. Specifically, we first introduce a dual variational autoencoder to represent a joint distribution of paired heterogeneous images. Then, in order to ensure the identity consistency of the generated paired heterogeneous images, we impose a distribution alignment in the latent space and a pairwise identity preserving in the image space. Moreover, the HFR network reduces the domain discrepancy by constraining the pairwise feature distances between the generated paired heterogeneous images. Extensive experiments on four HFR databases show that our method can significantly improve state-of-the-art results. When using the generated paired images for training, our method gains more than 18\% True Positive Rate improvements over the baseline model when False Positive Rate is at $10^{-5}$.
Chaoyou Fu, Xiang Wu 0001, Yibo Hu 0001, Huaibo Huang, Ran He 0001
NeurIPS5
2019 Wavelet Domain Generative Adversarial Network for Multi-scale Face Hallucination
Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan
Int. J. Comput. Vis.2
2019 Wasserstein CNN: Learning Invariant Features for NIR-VIS Face Recognition
abstract
Heterogeneous face recognition (HFR) aims at matching facial images acquired from different sensing modalities with mission-critical applications in forensics, security and commercial sectors. However, HFR presents more challenging issues than traditional face recognition because of the large intra-class variation among heterogeneous face images and the limited availability of training samples of cross-modality face image pairs. This paper proposes the novel Wasserstein convolutional neural network (WCNN) approach for learning invariant features between near-infrared (NIR) and visual (VIS) face images (i.e., NIR-VIS face recognition). The low-level layers of the WCNN are trained with widely available face images in the VIS spectrum, and the high-level layer is divided into three parts: the NIR layer, the VIS layer and the NIR-VIS shared layer. The first two layers aim at learning modality-specific features, and the NIR-VIS shared layer is designed to learn a modality-invariant feature subspace. The Wasserstein distance is introduced into the NIR-VIS shared layer to measure the dissimilarity between heterogeneous feature distributions. W-CNN learning is performed to minimize the Wasserstein distance between the NIR distribution and the VIS distribution for invariant deep feature representations of heterogeneous face images. To avoid the over-fitting problem on small-scale heterogeneous face data, a correlation prior is introduced on the fully-connected WCNN layers to reduce the size of the parameter space. This prior is implemented by a low-rank constraint in an end-to-end network. The joint formulation leads to an alternating minimization for deep feature representation at the training stage and an efficient computation for heterogeneous data at the testing stage. Extensive experiments using three challenging NIR-VIS face recognition databases demonstrate the superiority of the WCNN method over state-of-the-art methods.
Ran He 0001, Xiang Wu 0001, Zhenan Sun, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Aggregating Randomized Clustering-Promoting Invariant Projections for Domain Adaptation
abstract
Unsupervised domain adaptation aims to leverage the labeled source data to learn with the unlabeled target data. Previous trandusctive methods tackle it by iteratively seeking a low-dimensional projection to extract the invariant features and obtaining the pseudo target labels via building a classifier on source data. However, they merely concentrate on minimizing the cross-domain distribution divergence, while ignoring the intra-domain structure especially for the target domain. Even after projection, possible risk factors like imbalanced data distribution may still hinder the performance of target label inference. In this paper, we propose a simple yet effective domain-invariant projection ensemble approach to tackle these two issues together. Specifically, we seek the optimal projection via a novel relaxed domain-irrelevant clustering-promoting term that jointly bridges the cross-domain semantic gap and increases the intra-class compactness in both domains. To further enhance the target label inference, we first develop a 'sampling-and-fusion' framework, under which multiple projections are independently learned based on various randomized coupled domain subsets. Subsequently, aggregating models such as majority voting are utilized to leverage multiple projections and classify unlabeled target data. Extensive experimental results on six visual benchmarks including object, face, and digit images, demonstrate that the proposed methods gain remarkable margins over state-of-the-art unsupervised domain adaptation methods.
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Learning a bi-level adversarial network with global and local perception for makeup-invariant face verification
Yi Li 0018, Lingxiao Song, Xiang Wu 0001, Ran He 0001, Tieniu Tan
Pattern Recognit.4
2019 Exploring uncertainty in pseudo-label guided unsupervised domain adaptation
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan
Pattern Recognit.2
2019 3D Aided Duet GANs for Multi-View Face Image Synthesis
abstract
Multi-view face synthesis from a single image is an ill-posed computer vision problem. It often suffers from appearance distortions if it is not well-defined. Producing photo-realistic and identity preserving multi-view results is still a not well-defined synthesis problem. This paper proposes 3D aided duet generative adversarial networks (AD-GAN) to precisely rotate the yaw angle of an input face image to any specified angle. AD-GAN decomposes the challenging synthesis problem into two well-constrained subtasks that correspond to a face normalizer and a face editor. The normalizer first frontalizes an input image, and then the editor rotates the frontalized image to a desired pose guided by a remote code. In the meantime, the face normalizer is designed to estimate a novel dense UV correspondence field, making our model aware of 3D face geometry information. In order to generate photo-realistic local details and accelerate convergence process, the normalizer and the editor are trained in a two-stage manner and regulated by a conditional self-cycle loss and a perceptual loss. Exhaustive experiments on both controlled and uncontrolled environments demonstrate that the proposed method not only improves the visual realism of multi-view synthetic images but also preserves identity information well.
Jie Cao 0002, Yibo Hu 0001, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2019 Global and Local Consistent Wavelet-Domain Age Synthesis
abstract
Age synthesis is a challenging task due to the complicated and non-linear transformation in the human aging process. Aging information is usually reflected in local facial parts, such as wrinkles at the eye corners. However, these local facial parts contribute less in previous GAN-based methods for age synthesis. To address this issue, we propose a wavelet-domain global and local consistent age generative adversarial network (WaveletGLCA-GAN), in which one global specific network and three local specific networks are integrated together to capture both global topology information and local texture details of human faces. Different from the most existing methods that modeling age synthesis in image domain, we adopt wavelet transform to depict the textual information in frequency domain. Moreover, five types of losses are adopted: 1) adversarial loss aims to generate realistic wavelets; 2) identity preserving loss aims to better preserve identity information; 3) age preserving loss aims to enhance the accuracy of age synthesis; 4) pixel-wise loss aims to preserve the background information of the input face; and 5) the total variation regularization aims to remove ghosting artifacts. Our method is evaluated on three face aging datasets, including CACD2000, Morph, and FG-NET. Qualitative and quantitative experiments show the superiority of the proposed method over other state-of-the-arts.
Peipei Li 0002, Yibo Hu 0001, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.3
2018 Anti-Makeup: Learning A Bi-Level Adversarial Network for Makeup-Invariant Face Verification
abstract
Makeup is widely used to improve facial attractiveness and is well accepted by the public. However, different makeup styles will result in significant facial appearance changes. It remains a challenging problem to match makeup and non-makeup face images. This paper proposes a learning from generation approach for makeup-invariant face verification by introducing a bi-level adversarial network (BLAN). To alleviate the negative effects from makeup, we first generate non-makeup images from makeup ones, and then use the synthesized non-makeup images for further verification. Two adversarial networks in BLAN are integrated in an end-to-end deep network, with the one on pixel level for reconstructing appealing facial images and the other on feature level for preserving identity information. These two networks jointly reduce the sensing gap between makeup and non-makeup images. Moreover, we make the generator well constrained by incorporating multiple perceptual losses. Experimental results on three benchmark makeup face datasets demonstrate that our method achieves state-of-the-art verification accuracy across makeup status and can produce photo-realistic non-makeup face images.
Yi Li 0018, Lingxiao Song, Xiang Wu 0001, Ran He 0001, Tieniu Tan
AAAI4
2018 Adversarial Discriminative Heterogeneous Face Recognition
abstract
The gap between sensing patterns of different face modalities remains a challenging problem in heterogeneous face recognition (HFR). This paper proposes an adversarial discriminative feature learning framework to close the sensing gap via adversarial learning on both raw-pixel space and compact feature space. This framework integrates cross-spectral face hallucination and discriminative feature learning into an end-to-end adversarial network. In the pixel space, we make use of generative adversarial networks to perform cross-spectral face hallucination. An elaborate two-path model is introduced to alleviate the lack of paired images, which gives consideration to both global structures and local textures. In the feature space, an adversarial loss and a high-order variance discrepancy loss are employed to measure the global and local discrepancy between two heterogeneous distributions respectively. These two losses enhance domain-invariant feature learning and modality independent noise removing. Experimental results on three NIR-VIS databases show that our proposed approach outperforms state-of-the-art HFR methods, without requiring of complex network or large-scale training dataset.
Lingxiao Song, Man Zhang 0005, Xiang Wu 0001, Ran He 0001
AAAI4
2018 Coupled Deep Learning for Heterogeneous Face Recognition
abstract
Heterogeneous face matching is a challenge issue in face recognition due to large domain difference as well as insufficient pairwise images in different modalities during training. This paper proposes a coupled deep learning (CDL) approach for the heterogeneous face matching. CDL seeks a shared feature space in which the heterogeneous face matching problem can be approximately treated as a homogeneous face matching problem. The objective function of CDL mainly includes two parts. The first part contains a trace norm and a block-diagonal prior as relevance constraints, which not only make unpaired images from multiple modalities be clustered and correlated, but also regularize the parameters to alleviate overfitting. An approximate variational formulation is introduced to deal with the difficulties of optimizing low-rank constraint directly. The second part contains a cross modal ranking among triplet domain specific images to maximize the margin for different identities and increase data for a small amount of training samples. Besides, an alternating minimization method is employed to iteratively update the parameters of CDL. Experimental results show that CDL achieves better performance on the challenging CASIA NIR-VIS 2.0 face recognition database, the IIIT-D Sketch database, the CUHK Face Sketch (CUFS), and the CUHK Face Sketch FERET (CUFSF), which significantly outperforms state-of-the-art heterogeneous face recognition methods.
Xiang Wu 0001, Lingxiao Song, Ran He 0001, Tieniu Tan
AAAI3
2018 X-GACMN: An X-Shaped Generative Adversarial Cross-Modal Network with Hypersphere Embedding
Weikuo Guo, Jian Liang 0001, Xiangwei Kong 0001, Lingxiao Song, Ran He 0001
ACCV (5)5
2018 Pose-Guided Photorealistic Face Rotation
abstract
Face rotation provides an effective and cheap way for data augmentation and representation learning of face recognition. It is a challenging generative learning problem due to the large pose discrepancy between two face images. This work focuses on flexible face rotation of arbitrary head poses, including extreme profile views. We propose a novel Couple-Agent Pose-Guided Generative Adversarial Network (CAPG-GAN) to generate both neutral and profile head pose face images. The head pose information is encoded by facial landmark heatmaps. It not only forms a mask image to guide the generator in learning process but also provides a flexible controllable condition during inference. A couple-agent discriminator is introduced to reinforce on the realism of synthetic arbitrary view faces. Besides the generator and conditional adversarial loss, CAPG-GAN further employs identity preserving loss and total variation regularization to preserve identity information and refine local textures respectively. Quantitative and qualitative experimental results on the Multi-PIE and LFW databases consistently show the superiority of our face rotation method over the state-of-the-art.
Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun
CVPR4
2018 Learning Discriminative Geodesic Flow Kernel for Unsupervised Domain Adaptation
abstract
Extracting the domain-invariant features provides an important intuition for unsupervised domain adaptation. Due to the unavailable target labels, it is difficult to guarantee that the learned domain-invariant features are good for target instances classification. In this paper, we extend the classic geodesic flow kernel method by leveraging the pseudo labels during the training process to learn a discriminative geodesic flow kernel for unsupervised domain adaptation. Specifically, the proposed method alternately discovers the pseudo target labels and builds the geodesic flow from a discriminative source subspace to another ‘discriminative’ target subspace. More specially, the pseudo target labels are inferred via the learned kernel based on an easy yet effective label propagation strategy. Hence, the proposed method not only holds the property of domain-invariance, but also maximizes the consistency between pseudo label structure and data structure. Experimental results illustrate that the proposed method outperforms the state-of-the-art unsupervised domain adaptation methods for object recognition and sentiment analysis.
Jianze Wei, Jian Liang 0001, Ran He 0001, Jinfeng Yang
ICME3
2018 Global and Local Consistent Age Generative Adversarial Networks
abstract
Age progression/regression is a challenging task due to the complicated and non-linear transformation in human aging process. Many researches have shown that both global and local facial features are essential for face representation [1], but previous GAN based methods mainly focused on the global feature in age synthesis. To utilize both global and local facial information, we propose a Global and Local Consistent Age Generative Adversarial Network (GLCA-GAN). In our generator, a global network learns the whole facial structure and simulates the aging trend of the whole face, while three crucial facial patches are progressed or regressed by three local networks aiming at imitating subtle changes of crucial facial subregions. To preserve most of the details in age-attribute-irrelevant areas, our generator learns the residual face. Moreover, we employ an identity preserving loss to better preserve the identity information, as well as age preserving loss to enhance the accuracy of age synthesis. A pixel loss is also adopted to preserve detailed facial information of the input face. Our proposed method is evaluated on three face aging datasets, i.e., CACD dataset, Morph dataset and FG-NET dataset. Experimental results show appealing performance of the proposed method by comparing with the state-of-the-art.
Peipei Li 0002, Yibo Hu 0001, Qi Li 0005, Ran He 0001, Zhenan Sun
ICPR4
2018 Conditional Expression Synthesis with Face Parsing Transformation
abstract
Facial expression synthesis with various intensities is a challenging synthesis task due to large identity appearance variations and a paucity of efficient means for intensity measurement. This paper advances the expression synthesis domain by the introduction of a Couple-Agent Face Parsing based Generative Adversarial Network (CAFP-GAN) that unites the knowledge of facial semantic regions and controllable expression signals. Specially, we employ a face parsing map as a controllable condition to guide facial texture generation with a special expression, which can provide a semantic representation of every pixel of facial regions. Our method consists of two sub-networks: face parsing prediction network (FPPN) uses controllable labels (expression and intensity) to generate a face parsing map transformation that corresponds to the labels from the input neutral face, and facial expression synthesis network (FESN) makes the pretrained FPPN as a part of it to provide the face parsing map as a guidance for expression synthesis. To enhance the reality of results, couple-agent discriminators are served to distinguish fake-real pairs in both two sub-nets. Moreover, we only need the neutral face and the labels to synthesize the unknown expression with different intensities. Experimental results on three popular facial expression databases show that our method has the compelling ability on continuous expression synthesis.
Zhihe Lu, Tanhao Hu, Lingxiao Song, Zhaoxiang Zhang 0001, Ran He 0001
ACM Multimedia5
2018 Geometry Guided Adversarial Facial Expression Synthesis
abstract
Facial expression synthesis has drawn much attention in the field of computer graphics and pattern recognition. It has been widely used in face animation and recognition. However, it is still challenging due to the high-level semantic presence of large and non-linear face geometry variations. This paper proposes a Geometry-Guided Generative Adversarial Network (G2-GAN) for continuously-adjusting and identity-preserving facial expression synthesis. We employ facial geometry (fiducial points) as a controllable condition to guide facial texture synthesis with specific expression. A pair of generative adversarial subnetworks is jointly trained towards opposite tasks: expression removal and expression synthesis. The paired networks form a mapping cycle between neutral expression and arbitrary expressions, with which the proposed approach can be conducted among unpaired data. The proposed paired networks also facilitate other applications such as face transfer, expression interpolation and expression-invariant face recognition. Experimental results on several facial expression databases show that our method can generate compelling perceptual results on different expression editing tasks.
Lingxiao Song, Zhihe Lu, Ran He 0001, Zhenan Sun, Tieniu Tan
ACM Multimedia3
2018 Learning a High Fidelity Pose Invariant Model for High-resolution Face Frontalization
abstract
Face frontalization refers to the process of synthesizing the frontal view of a face from a given profile. Due to self-occlusion and appearance distortion in the wild, it is extremely challenging to recover faithful results and preserve texture details in a high-resolution. This paper proposes a High Fidelity Pose Invariant Model (HF-PIM) to produce photographic and identity-preserving results. HF-PIM frontalizes the profiles through a novel texture warping procedure and leverages a dense correspondence field to bind the 2D and 3D surface spaces. We decompose the prerequisite of warping into dense correspondence field estimation and facial texture map recovering, which are both well addressed by deep networks. Different from those reconstruction methods relying on 3D data, we also propose Adversarial Residual Dictionary Learning (ARDL) to supervise facial texture map recovering with only monocular images. Exhaustive experiments on both controlled and uncontrolled environments demonstrate that the proposed method not only boosts the performance of pose-invariant face recognition but also dramatically improves high-resolution frontalization appearances.
Jie Cao 0002, Yibo Hu 0001, Hongwen Zhang 0001, Ran He 0001, Zhenan Sun
NeurIPS4
2018 IntroVAE: Introspective Variational Autoencoders for Photographic Image Synthesis
abstract
We present a novel introspective variational autoencoder (IntroVAE) model for synthesizing high-resolution photographic images. IntroVAE is capable of self-evaluating the quality of its generated samples and improving itself accordingly. Its inference and generator models are jointly trained in an introspective way. On one hand, the generator is required to reconstruct the input images from the noisy outputs of the inference model as normal VAEs. On the other hand, the inference model is encouraged to classify between the generated and real samples while the generator tries to fool it as GANs. These two famous generative frameworks are integrated in a simple yet efficient single-stream architecture that can be trained in a single stage. IntroVAE preserves the advantages of VAEs, such as stable training and nice latent manifold. Unlike most other hybrid models of VAEs and GANs, IntroVAE requires no extra discriminators, because the inference model itself serves as a discriminator to distinguish between the generated and real samples. Experiments demonstrate that our method produces high-resolution photo-realistic images (e.g., CELEBA images at (1024^{2})), which are comparable to or better than the state-of-the-art GANs.
Huaibo Huang, Zhihang Li, Ran He 0001, Zhenan Sun, Tieniu Tan
NeurIPS3
2018 Learning structured ordinal measures for video based face recognition
Ran He 0001, Tieniu Tan, Larry Davis 0001, Zhenan Sun
Pattern Recognit.1
2018 Robust linear representation via exploiting structure prior
Dong Wang 0004, Ran He 0001, Liang Wang 0001, Tieniu Tan
Pattern Recognit.2
2018 A Light CNN for Deep Face Representation With Noisy Labels
abstract
The volume of convolutional neural network (CNN) models proposed for face recognition has been continuously growing larger to better fit the large amount of training data. When training data are obtained from the Internet, the labels are likely to be ambiguous and inaccurate. This paper presents a Light CNN framework to learn a compact embedding on the large-scale face data with massive noisy labels. First, we introduce a variation of maxout activation, called max-feature-map (MFM), into each convolutional layer of CNN. Different from maxout activation that uses many feature maps to linearly approximate an arbitrary convex activation function, MFM does so via a competitive relationship. MFM can not only separate noisy and informative signals but also play the role of feature selection between two feature maps. Second, three networks are carefully designed to obtain better performance, meanwhile, reducing the number of parameters and computational costs. Finally, a semantic bootstrapping method is proposed to make the prediction of the networks more consistent with noisy labels. Experimental results show that the proposed framework can utilize large-scale noisy data to learn a Light model that is efficient in computational costs and storage spaces. The learned single network with a 256-D representation achieves state-of-the-art results on various face benchmarks without fine-tuning.
Xiang Wu 0001, Ran He 0001, Zhenan Sun, Tieniu Tan
IEEE Trans. Inf. Forensics Secur.2
2018 DeMeshNet: Blind Face Inpainting for Deep MeshFace Verification
abstract
MeshFace photos have been widely used in many Chinese business organizations to protect ID face photos from being misused. The occlusions incurred by random meshes severely degenerate the performance of face verification systems, which raises the MeshFace verification problem between MeshFace and daily photos. Previous methods cast this problem as a typical low-level vision problem, i.e., blind inpainting. They recover perceptually pleasing clear ID photos from MeshFaces by enforcing pixel level similarity between the recovered ID images and the ground-truth clear ID images and then perform face verification on them. Essentially, face verification is conducted on a compact feature space rather than the image pixel space. Therefore, this paper argues that pixel level similarity and feature level similarity jointly offer the key to improve the verification performance. Based on this insight, we offer a novel feature oriented blind face inpainting framework. Specifically, we implement this by establishing a novel DeMeshNet, which consists of three parts. The first part addresses blind inpainting of the MeshFaces by implicitly exploiting extra supervision from the occlusion position to enforce pixel level similarity. The second part explicitly enforces a feature level similarity in the compact feature space, which can explore informative supervision from the feature space to produce better inpainting results for verification. The last part copes with face alignment within the net via a customized spatial transformer module when extracting deep facial features. All three parts are implemented within an end-to-end network that facilitates efficient optimization. Extensive experiments on two MeshFace data sets demonstrate the effectiveness of the proposed DeMeshNet as well as the insight of this paper.
Shu Zhang 0015, Ran He 0001, Zhenan Sun, Tieniu Tan
IEEE Trans. Inf. Forensics Secur.2
2017 Self-Paced Learning: An Implicit Regularization Perspective
abstract
Self-paced learning (SPL) mimics the cognitive mechanism of humans and animals that gradually learns from easy to hard samples. One key issue in SPL is to obtain better weighting strategy that is determined by the minimizer function. Existing methods usually pursue this by artificially designing the explicit form of SPL regularizer. In this paper, we study a group of new regularizer (named self-paced implicit regularizer) that is deduced from robust loss function. Based on the convex conjugacy theory, the minimizer function for self-paced implicit regularizer can be directly learned from the latent loss function, while the analytic form of the regularizer can be even unknown. A general framework (named SPL-IR) for SPL is developed accordingly. We demonstrate that the learning procedure of SPL-IR is associated with latent robust loss functions, thus can provide some theoretical insights for its working mechanism. We further analyze the relation between SPL-IR and half-quadratic optimization and provide a group of self-paced implicit regularizer. Finally, we implement SPL-IR to both supervised and unsupervised tasks, and experimental results corroborate our ideas and demonstrate the correctness and effectiveness of implicit regularizers.
Yanbo Fan, Ran He 0001, Jian Liang 0001, Bao-Gang Hu
AAAI2
2017 Learning Invariant Deep Representation for NIR-VIS Face Recognition
abstract
Visual versus near infrared (VIS-NIR) face recognition is still a challenging heterogeneous task due to large appearance difference between VIS and NIR modalities. This paper presents a deep convolutional network approach that uses only one network to map both NIR and VIS images to a compact Euclidean space. The low-level layers of this network are trained only on large-scale VIS data. Each convolutional layer is implemented by the simplest case of maxout operator. The high-level layer is divided into two orthogonal subspaces that contain modality-invariant identity information and modality-variant spectrum information respectively. Our joint formulation leads to an alternating minimization approach for deep representation at the training time and an efficient computation for heterogeneous data at the testing time. Experimental evaluations show that our method achieves 94% verification rate at FAR=0.1% on the challenging CASIA NIR-VIS 2.0 face recognition dataset. Compared with state-of-the-art methods, it reduces the error rate by 58% only with a compact 64-D representation.
Ran He 0001, Xiang Wu 0001, Zhenan Sun, Tieniu Tan
AAAI1
2017 Automatic image cropping with aesthetic map and gradient energy map
abstract
Image cropping is a fundamental task in image editing to enhance the aesthetic quality of images. In this paper, we propose an automatic image cropping technique based on aesthetic map and gradient energy map. Instead of utilizing aesthetic rules in previous methods, we learn the aesthetic map by a deep convolutional neural network with a large-scale dataset for aesthetic quality assessment. The aesthetic map can highlight the discriminative image regions for high (or low) aesthetic quality category. The gradient energy map presents edge spatial distribution of images and is developed to compute the simplicity of images. Then a composition model is learned with the aesthetic map and gradient energy map to evaluate the quality of composition for crops. Moreover, an aesthetic preservation model is developed to compute the aesthetic information remained in crops to avoid cropping out high aesthetic regions. Experiments show that our approach significantly outperforms state-of-the-art cropping methods.
Yueying Kao, Ran He 0001, Kaiqi Huang
ICASSP2
2017 Fast multi-view face alignment via multi-task auto-encoders
abstract
Face alignment is an important problem in computer vision. It is still an open problem due to the variations of facial attributes (e.g., head pose, facial expression, illumination variation). Many studies have shown that face alignment and facial attribute analysis are often correlated. This paper develops a two-stage multi-task Auto-encoders framework for fast face alignment by incorporating head pose information to handle large view variations. In the first and second stages, multi-task Auto-encoders are used to roughly locate and further refine facial landmark locations with related pose information, respectively. Besides, the shape constraint is naturally encoded into our two-stage face alignment framework to preserve facial structures. A coarse-to-fine strategy is adopted to refine the facial landmark results with the shape constraint. Furthermore, the computational cost of our method is much lower than its deep learning competitors. Experimental results on various challenging datasets show the effectiveness of the proposed method.
Qi Li 0005, Zhenan Sun, Ran He 0001
IJCB3
2017 Wavelet-SRNet: A Wavelet-Based CNN for Multi-scale Face Super Resolution
abstract
Most modern face super-resolution methods resort to convolutional neural networks (CNN) to infer highresolution (HR) face images. When dealing with very low resolution (LR) images, the performance of these CNN based methods greatly degrades. Meanwhile, these methods tend to produce over-smoothed outputs and miss some textural details. To address these challenges, this paper presents a wavelet-based CNN approach that can ultra-resolve a very low resolution face image of 16 × 16 or smaller pixelsize to its larger version of multiple scaling factors (2×, 4×, 8× and even 16×) in a unified framework. Different from conventional CNN methods directly inferring HR images, our approach firstly learns to predict the LR's corresponding series of HR's wavelet coefficients before reconstructing HR images from them. To capture both global topology information and local texture details of human faces, we present a flexible and extensible convolutional neural network with three types of loss: wavelet prediction loss, texture loss and full-image loss. Extensive experiments demonstrate that the proposed approach achieves more appealing results both quantitatively and qualitatively than state-ofthe- art super-resolution methods.
Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan
ICCV2
2017 Unsupervised feature selection with ordinal locality
abstract
Unsupervised feature selection has shown significant potential in distance-based clustering tasks. This paper proposes a novel triplet induced method. Firstly, a triplet-based loss function is introduced to enforce the selected feature groups to preserve ordinal locality of original data, which contributes to distance-based clustering tasks. Secondly, we simplify the orthogonal basis clustering by imposing an orthogonal constraint on the feature projection matrix. Consequently, a general framework for simultaneous feature selection and clustering is discussed. Thirdly, an alternating minimization algorithm is employed to efficiently optimize the proposed model together with rapid convergence. Extensive comparison experiments on several benchmark datasets well validate the encouraging gain in clustering from our proposed method.
Jun Guo 0008, Yanqing Guo, Xiangwei Kong 0001, Ran He 0001
ICME4
2017 Deep Supervised Discrete Hashing
abstract
With the rapid growth of image and video data on the web, hashing has been extensively studied for image or video search in recent years. Benefiting from recent advances in deep learning, deep hashing methods have achieved promising results for image retrieval. However, there are some limitations of previous deep hashing methods (e.g., the semantic information is not fully exploited). In this paper, we develop a deep supervised discrete hashing algorithm based on the assumption that the learned binary codes should be ideal for classification. Both the pairwise label information and the classification information are used to learn the hash codes within one stream framework. We constrain the outputs of the last layer to be binary codes directly, which is rarely investigated in deep hashing algorithm. Because of the discrete nature of hash codes, an alternating minimization method is used to optimize the objective function. Experimental results have shown that our method outperforms current state-of-the-art methods on benchmark datasets.
Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan
NIPS3
2017 Editorial: Special issue on ubiquitous biometrics
Ran He 0001, Brian C. Lovell, Rama Chellappa, Anil K. Jain 0001, Zhenan Sun
Pattern Recognit.1
2017 Deep Aesthetic Quality Assessment With Semantic Information
abstract
Human beings often assess the aesthetic quality of an image coupled with the identification of the image's semantic content. This paper addresses the correlation issue between automatic aesthetic quality assessment and semantic recognition. We cast the assessment problem as the main task among a multi-task deep model, and argue that semantic recognition task offers the key to address this problem. Based on convolutional neural networks, we employ a single and simple multi-task framework to efficiently utilize the supervision of aesthetic and semantic labels. A correlation item between these two tasks is further introduced to the framework by incorporating the inter-task relationship learning. This item not only provides some useful insight about the correlation but also improves assessment accuracy of the aesthetic task. In particular, an effective strategy is developed to keep a balance between the two tasks, which facilitates to optimize the parameters of the framework. Extensive experiments on the challenging Aesthetic Visual Analysis dataset and Photo.net dataset validate the importance of semantic recognition in aesthetic quality assessment, and demonstrate that multitask deep models can discover an effective aesthetic representation to achieve the state-of-the-art results.
Yueying Kao, Ran He 0001, Kaiqi Huang
IEEE Trans. Image Process.2
2017 Image Piece Learning for Weakly Supervised Semantic Segmentation
abstract
The task of semantic segmentation is to infer a predefined category label for each pixel in the image. For most cases, image segmentation is established as a fully supervised task. These methods all built on the basis of having access to sufficient pixel-wise annotated samples for training. However, obtaining the satisfied ground truth is not only labor intensive but also time-consuming, which severely hinders the generality of these fully supervised methods. Instead of pixel-level ground truth, weakly supervised approaches learn their models from much less prior information, e.g., image-level annotation. In this paper, we propose a novel conditional random field (CRF) based framework for weakly supervised semantic segmentation. Enlightened by jigsaw puzzles, we start the approach with merging superpixels from an image into larger pieces by a newly designed strategy. Then pieces from all the training images are gathered and associated with appropriate semantic labels by CRF. Thus, the piece library is constructed, achieving remarkable universality and flexibility. In the case of testing, we compare the superpixels with image pieces in the library and assign them the labels that minimize the potential energy. In addition, the proposed framework is fit for domain adaption and obtains promising results, which is of great practical value. Extensive experimental results on PASCAL VOC 2007, MSRC-21, and VOC 2012 databases demonstrate that our framework outperforms or is comparable to state-of-the-art segmentation methods.
Yi Li 0018, Yanqing Guo, Yueying Kao, Ran He 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2016 Discriminative Analysis Dictionary Learning
abstract
Dictionary learning (DL) has been successfully applied to various pattern classification tasks in recent years. However, analysis dictionary learning (ADL), as a major branch of DL, has not yet been fully exploited in classification due to its poor discriminability. This paper presents a novel DL method, namely Discriminative Analysis Dictionary Learning (DADL), to improve the classification performance of ADL. First, a code consistent term is integrated into the basic analysis model to improve discriminability. Second, a triplet constraint-based local topology preserving loss function is introduced to capture the discriminative geometrical structures embedded in data. Third, correntropy induced metric is employed as a robust measure to better control outliers for classification. Then, half-quadratic minimization and alternate search strategy are used to speed up the optimization process so that there exist closed-form solutions in each alternating minimization stage. Experiments on several commonly used databases show that our proposed method not only significantly improves the discriminative ability of ADL, but also outperforms state-of-the-art synthesis DL methods.
Jun Guo 0008, Yanqing Guo, Xiangwei Kong 0001, Man Zhang 0005, Ran He 0001
AAAI5
2016 Simultaneous Feature and Sample Reduction for Image-Set Classification
abstract
Image-set classification is the assignment of a label to a given image set. In real-life scenarios such as surveillance videos, each image set often contains much redundancy in terms of features and samples. This paper introduces a joint learning method for image-set classification that simultaneously learns compact binary codes and removes redundant samples. The joint objective function of our model mainly includes two parts. The first part seeks a hashing function to generate binary codes that have larger inter-class and smaller intra-class distances. The second one reduces redundant samples with discrete constraints in a low-rank way. A kernel method based on anchor points is further used to reduce sample variations. The proposed discrete objective function is simplified to a series of sub-problems that admit an analytical solution, resulting in a high-quality discrete solution with a low computational cost. Experiments on three commonly used image-set datasets show that the proposed method for the tasks of face recognition from image sets is efficient and effective.
Man Zhang 0005, Ran He 0001, Zhenan Sun, Tieniu Tan
AAAI2
2016 Localize heavily occluded human faces via deep segmentation
abstract
Localizing heavily occluded human faces is a challenging problem in facial detection. Previous methods mainly employ sliding windows by determining whether windows include human faces. In this paper, we provide a novel segmentation-based perspective for heavily occluded face localization with deep convolutional neural networks (CNN). Our model takes an image as input without complicated pre-processing. After several convolutional layers, fully-connected layers and a softmax classifier, we can predict the labels of all pixels in an image, which is the key to localize heavily occluded human faces. Finally, we search a minimal rectangle to localize the human face. Our detector needs neither complex pre-processing nor the time-consuming sliding window. Besides, we use a single model to localize faces to further alleviate computational complexity. Experimental results show that our proposed method is a very effective way to localize heavily occluded human face.
Kaihao Zhang, Yongzhen Huang, Ran He 0001, Liang Wang 0001
ICIP3
2016 Group-Invariant Cross-Modal Subspace Learning
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan
IJCAI2
2016 Locally imposing function for Generalized Constraint Neural Networks - A study on equality constraints
abstract
This work is a further study on the Generalized Constraint Neural Network (GCNN) model [1], [2]. Two challenges are encountered in the study, that is, to embed any type of prior information and to select its imposing schemes. The work focuses on the second challenge and studies a new constraint imposing scheme for equality constraints. A new method called locally imposing function (LIF) is proposed to provide a local correction to the GCNN prediction function, which therefore falls within Locally Imposing Scheme (LIS). In comparison, the conventional Lagrange multiplier method is considered as Globally Imposing Scheme (GIS) because its added constraint term exhibits a global impact to its objective function. Two advantages are gained from LIS over GIS. First, LIS enables constraints to fire locally and explicitly in the domain only where they need on the prediction function. Second, constraints can be implemented within a network setting directly. We attempt to interpret several constraint methods graphically from a viewpoint of the locality principle. Numerical examples confirm the advantages of the proposed method. In solving boundary value problems with Dirichlet and Neumann constraints, the GCNN model with LIF is possible to achieve an exact satisfaction of the constraints.
Linlin Cao, Ran He 0001, Bao-Gang Hu
IJCNN2
2016 Topology preserving dictionary learning for pattern classification
abstract
In recent years, dictionary learning (DL) has shown significant potential in various classification tasks. However, most of previous works aim to learn a synthesis dictionary. The other major category of DL-analysis dictionary learning has not been fully exploited yet. This paper proposes a novel DL method, named Topology Preserving Dictionary Learning (TPDL). First, we propose a triplet-constraint-based topology preserving loss function to capture the underlying local topological structures of data in a supervised manner. Second, a sparse-label-matrix-based function is integrated into the basic analysis model to improve discriminative ability. Third, Huber M-estimator is employed as a robust metric to handle the errors (e.g., outliers and noise) that possibly exist in data. Then, an alternating optimization algorithm is developed based on half-quadratic minimization and alternate search strategy. Closed-form solutions in each alternating optimization stage speed up the minimization process. Experiments on four commonly used datasets show that our proposed TPDL achieves competitive performance in contrast to state-of-the-art DL methods.
Jun Guo 0008, Yanqing Guo, Bo Wang 0024, Xiangwei Kong 0001, Ran He 0001
IJCNN5
2016 Discrete Cross-Modal Hashing for Efficient Multimedia Retrieval
abstract
Hashing techniques have been widely adopted for cross-modal retrieval due to its low storage cost and fast query speed. Most existing cross-modal hashing methods aim to map heterogeneous data into the common low-dimensional hamming space and then threshold to obtain binary codes by relaxing the discrete constraint. However, this independent relaxation step also brings quantization errors, resulting in poor retrieval performances. Other cross-modal hashing methods try to directly optimize the challenging objective function with discrete binary constraints. Inspired by [1], we propose a novel supervised cross-modal hashing method called Discrete Cross-Modal Hashing (DCMH) to learn the discrete binary codes without relaxing them. DCMH is formulated through reconstructing the semantic similarity matrix and learning binary codes as ideal features for classification. Furthermore, DCMH alternately updates binary codes of each modality, and iteratively learns the discrete hashing codes bit by bit efficiently, which is quite promising for large-scale datasets. Extensive empirical results on three real-world datasets show that DCMH outperforms the baseline approaches significantly.
Dekui Ma, Jian Liang 0001, Xiangwei Kong 0001, Ran He 0001, Ying Li 0016
ISM4
2016 Frustratingly Easy Cross-Modal Hashing
abstract
Cross-modal hashing has attracted considerable attention due to its low storage cost and fast retrieval speed. Recently, more and more sophisticated researches related to this topic are proposed. However, they seem to be inefficient computationally for several reasons. On one hand, learning coupled hash projections makes the iterative optimization problem challenging. On the other hand, individual collective binary codes for each content are also learned with a high computation complexity. In this paper we describe a simple yet effective cross-modal hashing approach that can be implemented in just three lines of code. This approach first obtains the binary codes for one modality via unimodal hashing methods (e.g., iterative quantization (ITQ)), then applies simple linear regression to project the other modalities into the obtained binary subspace. Obviously, it is non-iterative and parameter-free, which makes it more attractive for many real-world applications. We further compare our approach with other state-of-the-art methods on four benchmark datasets (i.e., the Wiki, VOC, LabelMe and NUS-WIDE datasets). Despite its extraordinary simplicity, our approach performs remarkably and generally well for these datasets under different experimental settings (i.e., large-scale, high-dimensional and multi-label datasets).
Dekui Ma, Jian Liang 0001, Xiangwei Kong 0001, Ran He 0001
ACM Multimedia4
2016 Self-Paced Cross-Modal Subspace Matching
abstract
Cross-modal matching methods match data from different modalities according to their similarities. Most existing methods utilize label information to reduce the semantic gap between different modalities. However, it is usually time-consuming to manually label large-scale data. This paper proposes a Self-Paced Cross-Modal Subspace Matching (SCSM) method for unsupervised multimodal data. We assume that multimodal data are pair-wised and from several semantic groups, which form hard pair-wised constraints and soft semantic group constraints respectively. Then, we formulate the unsupervised cross-modal matching problem as a non-convex joint feature learning and data grouping problem. Self-paced learning, which learns samples from 'easy' to 'complex', is further introduced to refine the grouping result. Moreover, a multimodal graph is constructed to preserve the relationship of both inter- and intra-modality similarity. An alternating minimization method is employed to minimize the non-convex optimization problem, followed by the discussion on its convergence analysis and computational complexity. Experimental results on four multimodal databases show that SCSM outperforms state-of-the-art cross-modal subspace learning methods.
Jian Liang 0001, Zhihang Li, Ran He 0001, Jingdong Wang 0001
SIGIR4
2016 Joint Feature Selection and Subspace Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval has recently drawn much attention due to the widespread existence of multimodal data. It takes one type of data as the query to retrieve relevant data objects of another type, and generally involves two basic problems: the measure of relevance and coupled feature selection. Most previous methods just focus on solving the first problem. In this paper, we aim to deal with both problems in a novel joint learning framework. To address the first problem, we learn projection matrices to map multimodal data into a common subspace, in which the similarity between different modalities of data can be measured. In the learning procedure, the l21-norm penalties are imposed on the projection matrices separately to solve the second problem, which selects relevant and discriminative features from different feature spaces simultaneously. A multimodal graph regularization term is further imposed on the projected data,which preserves the inter-modality and intra-modality similarity relationships.An iterative algorithm is presented to solve the proposed joint learning problem, along with its convergence analysis. Experimental results on cross-modal retrieval tasks demonstrate that the proposed method outperforms the state-of-the-art subspace approaches.
Kaiye Wang, Ran He 0001, Liang Wang 0001, Wei Wang 0115, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Transformation invariant subspace clustering
Qi Li 0005, Zhenan Sun, Zhouchen Lin, Ran He 0001, Tieniu Tan
Pattern Recognit.4
2016 Information Theoretic Subspace Clustering
abstract
This paper addresses the problem of grouping the data points sampled from a union of multiple subspaces in the presence of outliers. Information theoretic objective functions are proposed to combine structured low-rank representations (LRRs) to capture the global structure of data and information theoretic measures to handle outliers. In theoretical part, we point out that group sparsity-induced measures (ℓ2,1-norm, ℓα-norm, and correntropy) can be justified from the viewpoint of halfquadratic (HQ) optimization, which facilitates both convergence study and algorithmic development. In particular, a general formulation is accordingly proposed to unify HQ-based group sparsity methods into a common framework. In algorithmic part, we develop information theoretic subspace clustering methods via correntropy. With the help of Parzen window estimation, correntropy is used to handle either outliers under any distributions or sample-specific errors in data. Pairwise link constraints are further treated as a prior structure of LRRs. Based on the HQ framework, iterative algorithms are developed to solve the nonconvex information theoretic loss functions. Experimental results on three benchmark databases show that our methods can further improve the robustness of LRR subspace clustering and outperform other state-of-the-art subspace clustering methods.
Ran He 0001, Liang Wang 0001, Zhenan Sun, Yingya Zhang, Bo Li 0005
IEEE Trans. Neural Networks Learn. Syst.1
2015 Multi-view Clustering via Structured Low-rank Representation
abstract
In this paper, we present a novel solution to multi-view clustering through a structured low-rank representation. When assuming similar samples can be linearly reconstructed by each other, the resulting representational matrix reflects the cluster structure and should ideally be block diagonal. We first impose low-rank constraint on the representational matrix to encourage better grouping effect. Then representational matrices under different views are allowed to communicate with each other and share their mutual cluster structure information. We develop an effective algorithm inspired by iterative re-weighted least squares for solving our formulation. During the optimization process, the intermediate representational matrix from one view serves as a cluster structure constraint for that from another view. Such mutual structural constraint fine-tunes the cluster structures from both views and makes them more and more agreeable. Extensive empirical study manifests the superiority and efficacy of the proposed method.
Dong Wang 0004, Qiyue Yin, Ran He 0001, Liang Wang 0001, Tieniu Tan
CIKM3
2015 A Two-step Approach to Cross-modal Hashing
abstract
With the rapid growth of multimedia data, it is very desirable to effectively and efficiently search objects of interest across different modalities from large scale databases. Cross-modal hashing provides a very promising way to address such problem. In this paper, we propose a two-step cross-modal hashing approach to obtain compact hash codes and learn hash functions from multimodal data. Our approach decomposes the cross-modal hashing problem into two steps: generating hash code and learning hash function. In the first step, we obtain the hash codes for all modalities of data via a joint multi-modal graph, which takes into consideration both the intra-modality and inter-modality similarity. In the second step, learning hashing function is formulated as a binary classification problem. We train binary classifiers to predict the hash code for any data object unseen before. Experimental results on two cross-modal datasets show the effectiveness of our proposed approach.
Kaiye Wang, Wei Wang 0115, Liang Wang 0001, Ran He 0001
ICMR4
2015 Multi-view clustering via pairwise sparse subspace representation
Qiyue Yin, Ran He 0001, Liang Wang 0001
Neurocomputing3
2015 Learning predictable binary codes for face indexing
Ran He 0001, Yinghao Cai, Tieniu Tan, Larry Davis 0001
Pattern Recognit.1
2015 Code Consistent Hashing Based on Information-Theoretic Criterion
abstract
Learning based hashing techniques have attracted broad research interests in the Big Media research area. They aim to learn compact binary codes which can preserve semantic similarity in the Hamming embedding. However, the discrete constraints imposed on binary codes typically make hashing optimizations very challenging. In this paper, we present a code consistent hashing (CCH) algorithm to learn discrete binary hash codes. To form a simple yet efficient hashing objective function, we introduce a new code consistency constraint to leverage discriminative information and propose to utilize the Hadamard code which favors an information-theoretic criterion as the class prototype. By keeping the discrete constraint and introducing an orthogonal constraint, our objective function can be minimized efficiently. Experimental results on three benchmark datasets demonstrate that the proposed CCH outperforms state-of-the-art hashing methods in both image retrieval and classification tasks, especially with short binary codes.
Shu Zhang 0015, Jian Liang 0001, Ran He 0001, Zhenan Sun
IEEE Trans. Big Data3
2015 Robust Subspace Clustering With Complex Noise
abstract
Subspace clustering has important and wide applications in computer vision and pattern recognition. It is a challenging task to learn low-dimensional subspace structures due to complex noise existing in high-dimensional data. Complex noise has much more complex statistical structures, and is neither Gaussian nor Laplacian noise. Recent subspace clustering methods usually assume a sparse representation of the errors incurred by noise and correct these errors iteratively. However, large corruptions incurred by complex noise cannot be well addressed by these methods. A novel optimization model for robust subspace clustering is proposed in this paper. Its objective function mainly includes two parts. The first part aims to achieve a sparse representation of each high-dimensional data point with other data points. The second part aims to maximize the correntropy between a given data point and its low-dimensional representation with other points. Correntropy is a robust measure so that the influence of large corruptions on subspace clustering can be greatly suppressed. An extension of pairwise link constraints is also proposed as prior information to deal with complex noise. Half-quadratic minimization is provided as an efficient solution to the proposed robust subspace clustering formulations. Experimental results on three commonly used data sets show that our method outperforms state-of-the-art subspace clustering methods.
Ran He 0001, Yingya Zhang, Zhenan Sun, Qiyue Yin
IEEE Trans. Image Process.1
2015 Cross-Modal Subspace Learning via Pairwise Constraints
abstract
In multimedia applications, the text and image components in a web document form a pairwise constraint that potentially indicates the same semantic concept. This paper studies cross-modal learning via the pairwise constraint and aims to find the common structure hidden in different modalities. We first propose a compound regularization framework to address the pairwise constraint, which can be used as a general platform for developing cross-modal algorithms. For unsupervised learning, we propose a multi-modal subspace clustering method to learn a common structure for different modalities. For supervised learning, to reduce the semantic gap and the outliers in pairwise constraints, we propose a cross-modal matching method based on compound ℓ21 regularization. Extensive experiments demonstrate the benefits of joint text and image modeling with semantically induced pairwise constraints, and they show that the proposed cross-modal methods can further reduce the semantic gap between different modalities and improve the clustering/matching accuracy.
Ran He 0001, Man Zhang 0005, Liang Wang 0001, Qiyue Yin
IEEE Trans. Image Process.1
2014 Jointly Learning Dictionaries and Subspace Structure for Video-Based Face Recognition
Guangxiao Zhang, Ran He 0001, Larry Davis 0001
ACCV (3)2
2014 Semi-supervised subspace segmentation
abstract
Subspace segmentation methods usually rely on the raw explicit feature vectors in an unsupervised manner. In many applications, it is cheap to obtain some pairwise link information that tells whether two data points are in the same subspace or not. Though partially available, such link information serves as some kind of high-level semantics, which can be further used as a constraint to improve the segmentation accuracy. By constructing a link matrix and using it as a regularizer, we propose a semi-supervised subspace segmentation model where the partially observed subspace membership prior can be encoded. Specificly, under the common linear representation assumption, we enforce the representational coefficient to be consistent with the link matrix. Thus the low-level and high-level information about the data can be integrated to produce more precise segmentation results. We then develop an effective algorithm to optimize our model in an alternating minimization way. Experimental results for both motion segmentation and face clustering validate that incorporating such link information is helpful to assist and bias the unsupervised subspace segmentation methods.
Dong Wang 0004, Qiyue Yin, Ran He 0001, Liang Wang 0001, Tieniu Tan
ICIP3
2014 Transform-invariant dictionary learning for face recognition
abstract
Dictionary learning has important applications in face recognition. However, large transformation variations of face images pose a grand challenge to conventional dictionary learning methods. A large portion of misleading dictionary atoms are usually learned to represent transformation factors, which will cause ambiguity in face recognition. To address this problem, this paper proposes a general framework for transform-invariant basis matrix learning. Specifically, we present a transform-invariant dictionary learning method which explicitly incorporates an appearance consistent error term to the original objective function in dictionary learning. The unified objective function is effectively optimized in an alternating iterative way. An ensemble of aligned images and a discriminative transform-invariant dictionary for sparse coding can be obtained by solving the formulated objective function. Experimental results on two public face databases demonstrate our algorithm's superiority compared with two state-of-the-art dictionary learning methods and the recently proposed transform-invariant PCA method.
Shu Zhang 0015, Man Zhang 0005, Ran He 0001, Zhenan Sun
ICIP3
2014 Robust Recovery of Corrupted Low-RankMatrix by Implicit Regularizers
abstract
Low-rank matrix recovery algorithms aim to recover a corrupted low-rank matrix with sparse errors. However, corrupted errors may not be sparse in real-world problems and the relationship between ℓ1 regularizer on noise and robust M-estimators is still unknown. This paper proposes a general robust framework for low-rank matrix recovery via implicit regularizers of robust M-estimators, which are derived from convex conjugacy and can be used to model arbitrarily corrupted errors. Based on the additive form of half-quadratic optimization, proximity operators of implicit regularizers are developed such that both low-rank structure and corrupted errors can be alternately recovered. In particular, the dual relationship between the absolute function in ℓ1 regularizer and Huber M-estimator is studied, which establishes a connection between robust low-rank matrix recovery methods and M-estimators based robust principal component analysis methods. Extensive experiments on synthetic and real-world data sets corroborate our claims and verify the robustness of the proposed framework.
Ran He 0001, Tieniu Tan, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Half-Quadratic-Based Iterative Minimization for Robust Sparse Representation
abstract
Robust sparse representation has shown significant potential in solving challenging problems in computer vision such as biometrics and visual surveillance. Although several robust sparse models have been proposed and promising results have been obtained, they are either for error correction or for error detection, and learning a general framework that systematically unifies these two aspects and explores their relation is still an open problem. In this paper, we develop a half-quadratic (HQ) framework to solve the robust sparse representation problem. By defining different kinds of half-quadratic functions, the proposed HQ framework is applicable to performing both error correction and error detection. More specifically, by using the additive form of HQ, we propose an ℓ1-regularized error correction method by iteratively recovering corrupted data from errors incurred by noises and outliers; by using the multiplicative form of HQ, we propose an ℓ1-regularized error detection method by learning from uncorrupted data iteratively. We also show that the ℓ1-regularization solved by soft-thresholding function has a dual relationship to Huber M-estimator, which theoretically guarantees the performance of robust sparse representation in terms of M-estimation. Experiments on robust face recognition under severe occlusion and corruption validate our framework and findings.
Ran He 0001, Wei-Shi Zheng 0001, Tieniu Tan, Zhenan Sun
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Gabor Ordinal Measures for Face Recognition
abstract
Great progress has been achieved in face recognition in the last three decades. However, it is still challenging to characterize the identity related features in face images. This paper proposes a novel facial feature extraction method named Gabor ordinal measures (GOM), which integrates the distinctiveness of Gabor features and the robustness of ordinal measures as a promising solution to jointly handle inter-person similarity and intra-person variations in face images. In the proposal, different kinds of ordinal measures are derived from magnitude, phase, real, and imaginary components of Gabor images, respectively, and then are jointly encoded as visual primitives in local regions. The statistical distributions of these visual primitives in face image blocks are concatenated into a feature vector and linear discriminant analysis is further used to obtain a compact and discriminative feature representation. Finally, a two-stage cascade learning method and a greedy block selection method are used to train a strong classifier for face recognition. Extensive experiments on publicly available face image databases, such as FERET, AR, and large scale FRGC v2.0, demonstrate state-of-the-art face recognition performance of GOM.
Zhenhua Chai, Zhenan Sun, Heydi Mendez Vazquez, Ran He 0001, Tieniu Tan
IEEE Trans. Inf. Forensics Secur.4
2013 Learning Coupled Feature Spaces for Cross-Modal Matching
abstract
Cross-modal matching has recently drawn much attention due to the widespread existence of multimodal data. It aims to match data from different modalities, and generally involves two basic problems: the measure of relevance and coupled feature selection. Most previous works mainly focus on solving the first problem. In this paper, we propose a novel coupled linear regression framework to deal with both problems. Our method learns two projection matrices to map multimodal data into a common feature space, in which cross-modal data matching can be performed. And in the learning procedure, the ell_21-norm penalties are imposed on the two projection matrices separately, which leads to select relevant and discriminative features from coupled feature spaces simultaneously. A trace norm is further imposed on the projected data as a low-rank constraint, which enhances the relevance of different modal data with connections. We also present an iterative algorithm based on half-quadratic minimization to solve the proposed regularized linear regression problem. The experimental results on two challenging cross-modal datasets demonstrate that the proposed method outperforms the state-of-the-art approaches.
Kaiye Wang, Ran He 0001, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICCV2
2013 Robust Subspace Clustering via Half-Quadratic Minimization
abstract
Subspace clustering has important and wide applications in computer vision and pattern recognition. It is a challenging task to learn low-dimensional subspace structures due to the possible errors (e.g., noise and corruptions) existing in high-dimensional data. Recent subspace clustering methods usually assume a sparse representation of corrupted errors and correct the errors iteratively. However large corruptions in real-world applications can not be well addressed by these methods. A novel optimization model for robust subspace clustering is proposed in this paper. The objective function of our model mainly includes two parts. The first part aims to achieve a sparse representation of each high-dimensional data point with other data points. The second part aims to maximize the correntropy between a given data point and its low-dimensional representation with other points. Correntropy is a robust measure so that the influence of large corruptions on subspace clustering can be greatly suppressed. An extension of our method with explicit introduction of representation error terms into the model is also proposed. Half-quadratic minimization is provided as an efficient solution to the proposed robust subspace clustering formulations. Experimental results on Hopkins 155 dataset and Extended Yale Database B demonstrate that our method outperforms state-of-the-art subspace clustering methods.
Yingya Zhang, Zhenan Sun, Ran He 0001, Tieniu Tan
ICCV3
2013 Robust spectral regression for face recognition
Yanqing Guo, Ran He 0001, Wei-Shi Zheng 0001, Xiangwei Kong 0001, Zhaofeng He 0001
Neurocomputing2
2013 A fast convex conjugated algorithm for sparse recovery
Ran He 0001, Xiao-Tong Yuan, Wei-Shi Zheng 0001
Neurocomputing1
2013 Two-Stage Nonnegative Sparse Representation for Large-Scale Face Recognition
abstract
This paper proposes a novel nonnegative sparse representation approach, called two-stage sparse representation (TSR), for robust face recognition on a large-scale database. Based on the divide and conquer strategy, TSR decomposes the procedure of robust face recognition into outlier detection stage and recognition stage. In the first stage, we propose a general multisubspace framework to learn a robust metric in which noise and outliers in image pixels are detected. Potential loss functions, including L1 , L2,1, and correntropy are studied. In the second stage, based on the learned metric and collaborative representation, we propose an efficient nonnegative sparse representation algorithm to find an approximation solution of sparse representation. According to the L1 ball theory in sparse representation, the approximated solution is unique and can be optimized efficiently. Then a filtering strategy is developed to avoid the computation of the sparse representation on the whole large-scale dataset. Moreover, theoretical analysis also gives the necessary condition for nonnegative least squares technique to find a sparse solution. Extensive experiments on several public databases have demonstrated that the proposed TSR approach, in general, achieves better classification accuracy than the state-of-the-art sparse representation methods. More importantly, a significant reduction of computational costs is reached in comparison with sparse representation classifier; this enables the TSR to be more suitable for robust face recognition on a large-scale dataset.
Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu, Xiangwei Kong 0001
IEEE Trans. Neural Networks Learn. Syst.1
2012 Semantic Pixel Sets Based Local Binary Patterns for Face Recognition
Zhenhua Chai, Heydi Mendez Vazquez, Ran He 0001, Zhenan Sun, Tieniu Tan
ACCV (2)3
2012 l2, 1 Regularized correntropy for robust feature selection
abstract
In this paper, we study the problem of robust feature extraction based on l2,1regularized correntropy in both theoretical and algorithmic manner. In theoretical part, we point out that an l2,1-norm minimization can be justified from the viewpoint of half-quadratic (HQ) optimization, which facilitates convergence study and algorithmic development. In particular, a general formulation is accordingly proposed to unify l1-norm and l2,1-norm minimization within a common framework. In algorithmic part, we propose an l2,1regularized correntropy algorithm to extract informative features meanwhile to remove outliers from training data. A new alternate minimization algorithm is also developed to optimize the non-convex correntropy objective. In terms of face recognition, we apply the proposed method to obtain an appearance-based model, called Sparse-Fisherfaces. Extensive experiments show that our method can select robust and sparse features, and outperforms several state-of-the-art subspace methods on largescale and open face recognition datasets.
Ran He 0001, Tieniu Tan, Liang Wang 0001, Wei-Shi Zheng 0001
CVPR1
2012 Robust large margin discriminant tangent analysis for face recognition
Nanhai Yang, Ran He 0001, Wei-Shi Zheng 0001, Xiukun Wang
Neural Comput. Appl.2
2012 Extracting non-negative basis images using pixel dispersion penalty
Wei-Shi Zheng 0001, Jian-Huang Lai, Shengcai Liao, Ran He 0001
Pattern Recognit.4
2012 Agglomerative Mean-Shift Clustering
abstract
Mean-Shift (MS) is a powerful nonparametric clustering method. Although good accuracy can be achieved, its computational cost is particularly expensive even on moderate data sets. In this paper, for the purpose of algorithmic speedup, we develop an agglomerative MS clustering method along with its performance analysis. Our method, namely Agglo-MS, is built upon an iterative query set compression mechanism which is motivated by the quadratic bounding optimization nature of MS algorithm. The whole framework can be efficiently implemented in linear running time complexity. We then extend Agglo-MS into an incremental version which performs comparably to its batch counterpart. The efficiency and accuracy of Agglo-MS are demonstrated by extensive comparing experiments on synthetic and real data sets.
Xiao-Tong Yuan, Bao-Gang Hu, Ran He 0001
IEEE Trans. Knowl. Data Eng.3
2011 Recovery of corrupted low-rank matrices via half-quadratic based nonconvex minimization
abstract
Recovering arbitrarily corrupted low-rank matrices arises in computer vision applications, including bioinformatic data analysis and visual tracking. The methods used involve minimizing a combination of nuclear norm and l1norm. We show that by replacing the l1norm on error items with nonconvex M-estimators, exact recovery of densely corrupted low-rank matrices is possible. The robustness of the proposed method is guaranteed by the M-estimator theory. The multiplicative form of half-quadratic optimization is used to simplify the nonconvex optimization problem so that it can be efficiently solved by iterative regularization scheme. Simulation results corroborate our claims and demonstrate the efficiency of our proposed method under tough conditions.
Ran He 0001, Zhenan Sun, Tieniu Tan, Wei-Shi Zheng 0001
CVPR1
2011 Nonnegative sparse coding for discriminative semi-supervised learning
abstract
An informative and discriminative graph plays an important role in the graph-based semi-supervised learning methods. This paper introduces a nonnegative sparse algorithm and its approximated algorithm based on the l0-l1equivalence theory to compute the nonnegative sparse weights of a graph. Hence, the sparse probability graph (SPG) is termed for representing the proposed method. The nonnegative sparse weights in the graph naturally serve as clustering indicators, benefiting for semi-supervised learning. More important, our approximation algorithm speeds up the computation of the nonnegative sparse coding, which is still a bottle-neck for any previous attempts of sparse non-negative graph learning. And it is much more efficient than using l1-norm sparsity technique for learning large scale sparse graph. Finally, for discriminative semi-supervised learning, an adaptive label propagation algorithm is also proposed to iteratively predict the labels of data on the SPG. Promising experimental results show that the nonnegative sparse coding is efficient and effective for discriminative semi-supervised learning.
Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu, Xiangwei Kong 0001
CVPR1
2011 Robust view transformation model for gait recognition
abstract
Recent gait recognition systems often suffer from the challenges including viewing angle variation and large intra-class variations. In order to address these challenges, this paper presents a robust View Transformation Model for gait recognition. Based on the gait energy image, the proposed method establishes a robust view transformation model via robust principal component analysis. Partial least square is used as feature selection method. Compared with the existing methods, the proposed method finds out a shared linear correlated low rank subspace, which brings the advantages that the view transformation model is robust to viewing angle variation, clothing and carrying condition changes. Conducted on the CASIA gait dataset, experimental results show that the proposed method outperforms the other existing methods.
Shuai Zheng 0001, Junge Zhang, Kaiqi Huang, Ran He 0001, Tieniu Tan
ICIP4
2011 A Regularized Correntropy Framework for Robust Pattern Recognition
abstract
This letter proposes a new multiple linear regression model using regularized correntropy for robust pattern recognition. First, we motivate the use of correntropy to improve the robustness of the classical mean square error (MSE) criterion that is sensitive to outliers. Then an l1regularization scheme is imposed on the correntropy to learn robust and sparse representations. Based on the half-quadratic optimization technique, we propose a novel algorithm to solve the nonlinear optimization problem. Second, we develop a new correntropy-based classifier based on the learned regularization scheme for robust object recognition. Extensive experiments over several applications confirm that the correntropy-based l1regularization can improve recognition accuracy and receiver operator characteristic curves under noise corruption and occlusion.
Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu, Xiangwei Kong 0001
Neural Comput.1
2011 Maximum Correntropy Criterion for Robust Face Recognition
abstract
In this paper, we present a sparse correntropy framework for computing robust sparse representations of face images for recognition. Compared with the state-of-the-art l(1)norm-based sparse representation classifier (SRC), which assumes that noise also has a sparse representation, our sparse algorithm is developed based on the maximum correntropy criterion, which is much more insensitive to outliers. In order to develop a more tractable and practical approach, we in particular impose nonnegativity constraint on the variables in the maximum correntropy criterion and develop a half-quadratic optimization technique to approximately maximize the objective function in an alternating way so that the complex optimization problem is reduced to learning a sparse representation through a weighted linear least squares problem with nonnegativity constraint at each iteration. Our extensive experiments demonstrate that the proposed method is more robust and efficient in dealing with the occlusion and corruption problems in face recognition as compared to the related state-of-the-art methods. In particular, it shows that the proposed method can improve both recognition accuracy and receiver operator characteristic (ROC) curves, while the computational cost is much lower than the SRC algorithms.
Ran He 0001, Wei-Shi Zheng 0001, Bao-Gang Hu
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Robust Principal Component Analysis Based on Maximum Correntropy Criterion
abstract
Principal component analysis (PCA) minimizes the mean square error (MSE) and is sensitive to outliers. In this paper, we present a new rotational-invariant PCA based on maximum correntropy criterion (MCC). A half-quadratic optimization algorithm is adopted to compute the correntropy objective. At each iteration, the complex optimization problem is reduced to a quadratic problem that can be efficiently solved by a standard optimization method. The proposed method exhibits the following benefits: 1) it is robust to outliers through the mechanism of MCC which can be more theoretically solid than a heuristic rule based on MSE; 2) it requires no assumption about the zero-mean of data for processing and can estimate data mean during optimization; and 3) its optimal solution consists of principal eigenvectors of a robust covariance matrix corresponding to the largest eigenvalues. In addition, kernel techniques are further introduced in the proposed method to deal with nonlinearly distributed data. Numerical results demonstrate that the proposed method can outperform robust rotational-invariant PCAs based on L(1) norm when outliers occur.
Ran He 0001, Bao-Gang Hu, Wei-Shi Zheng 0001, Xiangwei Kong 0001
IEEE Trans. Image Process.1
2010 Two-Stage Sparse Representation for Robust Recognition on Large-Scale Database
abstract
This paper proposes a novel robust sparse representation method, called the two-stage sparse representation (TSR), for robust recognition on a large-scale database. Based on the divide and conquer strategy, TSR divides the procedure of robust recognition into outlier detection stage and recognition stage. In the first stage, a weighted linear regression is used to learn a metric in which noise and outliers in image pixels are detected. In the second stage, based on the learnt metric, the large-scale dataset is firstly filtered into a small set according to the nearest neighbor criterion. Then a sparse representation is computed by the non-negative least squares technique. The sparse solution is unique and can be optimized efficiently. The extensive numerical experiments on several public databases demonstrate that the proposed TSR approach generally obtains better classification accuracy than the state of the art Sparse Representation Classification (SRC). At the same time, by using the TSR, a significant reduction of computational cost is reached by over fifty times in comparison with the SRC, which enables the TSR to be deployed more suitably for large-scale dataset.
Ran He 0001, Bao-Gang Hu, Wei-Shi Zheng 0001, Yanqing Guo
AAAI1
2010 Principal component analysis based on non-parametric maximum entropy
Ran He 0001, Bao-Gang Hu, Xiao-Tong Yuan, Wei-Shi Zheng 0001
Neurocomputing1
2009 Robust Discriminant Analysis Based on Nonparametric Maximum Entropy
Ran He 0001, Bao-Gang Hu, Xiao-Tong Yuan
ACML1
2009 Agglomerative Mean-Shift Clustering via Query Set Compression
abstract
Mean-Shift (MS) is a powerful non-parametric clustering method. Although good accuracy can be achieved, its computational cost is particularly expensive even on moderate data sets. In this paper, for the purpose of algorithm speedup, we develop an agglomerative MS clustering method called Agglo-MS, along with its mode-seeking ability and convergence property analysis. Our method is built upon an iterative query set compression mechanism which is motivated by the quadratic bounding optimization nature of MS. The whole framework can be efficiently implemented in linear running time complexity. Furthermore, we show that the pairwise constraint information can be naturally integrated into our framework to derive a semi-supervised non-parametric clustering method. Extensive experiments on toy and real-world data sets validate the speedup advantage and numerical accuracy of our method, as well as the superiority of its semi-supervised version.
Xiao-Tong Yuan, Bao-Gang Hu, Ran He 0001
SDM3
2008 Face shape recovery from a single image using CCA mapping between tensor spaces
abstract
In this paper, we propose a new approach for face shape recovery from a single image. A single near infrared (NIR) image is used as the input, and a mapping from the NIR tensor space to 3D tensor space, learned by using statistical learning, is used for the shape recovery. In the learning phase, the two tensor models are constructed for NIR and 3D images respectively, and a canonical correlation analysis (CCA) based multi-variate mapping from NIR to 3D faces is learned from a given training set of NIR-3D face pairs. In the reconstruction phase, given an NIR face image, the depth map is computed directly using the learned mapping with the help of tensor models. Experimental results are provided to evaluate the accuracy and speed of the method. The work provides a practical solution for reliable and fast shape recovery and modeling of 3D objects.
Zhen Lei 0001, Qinqun Bai, Ran He 0001, Stan Z. Li
CVPR3
2008 Regularized active shape model for shape alignment
abstract
Active shape model (ASM) statistically represents a shape by a set of well-defined landmark points and models object variations using principal component analysis (PCA). However, the extracted shape contour modeled by PCA is still unsmooth when the shape has a large variation compared with the mean shape. In this paper, we propose a regularized ASM (R-ASM) model for shape alignment. During training stage, we present a regularized shape subspace on which image smoothness constraint is imposed, such that the learned components to model shape variations should not only minimize reconstruction error but also obey smoothness principle. During searching stage, a coarse-to-fine parameter adjustment strategy is performed under Bayesian inference. It makes a desired shape smoother and more robust to local noise. Lastly, an inner shape is introduced to further regularize search results. Experiments on face alignment demonstrate the efficiency and effectiveness of our proposed approach.
Ran He 0001, Zhen Lei 0001, Xiao-Tong Yuan, Stan Z. Li
FG1
2008 Gabor volume based local binary pattern for face representation and recognition
abstract
This paper presents a novel face representation and recognition approach. The face image is first decomposed by multi-scale and multi-orientation Gabor filters and local binary pattern (LBP) analysis is then applied on the derived Gabor magnitude responses. Different from (W.C. Zhang et al., 2005), the present method not only describes the neighboring relationship in spatial domain, but also exploit those between different scales (frequency) and orientations. Specifically, we first reformulate the Gabor magnitude responses as a 3rd-order volume and then apply LBP analysis on three orthogonal planes of the Gabor volume, named GV-LBP-TOP in short, in a hope to encode sufficient information for face representation. Further, a computationally effective version, E-GV-LBP, is proposed to depict the neighboring changes in spatial, frequency and orientation domains simultaneously. Finally, the weighted histogram intersection metric is utilized to measure the dissimilarity of faces. Experimental results on FERET and FRGC ver 2.0 databases show the significant advantages of the proposed method.
Zhen Lei 0001, Shengcai Liao, Ran He 0001, Matti Pietikäinen, Stan Z. Li
FG3
2007 Learning Gabor Magnitude Features for Palmprint Recognition
Rufeng Chu, Zhen Lei 0001, Ran He 0001, Stan Z. Li
ACCV (2)4
2007 Coarse-to-Fine Statistical Shape Model by Bayesian Inference
Ran He 0001, Stan Z. Li, Zhen Lei 0001, Shengcai Liao
ACCV (1)1
2007 Color Constancy Via Convex Kernel Optimization
Xiao-Tong Yuan, Stan Z. Li, Ran He 0001
ACCV (1)3