VLDB 2026 Research / reviewers in the wild / expert
Huaibo Huang
dblp:211/7251
· DBLP profile ↗
74ranked-venue papers
7as first author
59since 2021 · last 2026
0000-0001-5866-2283ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 7 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 2 first-author · 33 since 2021Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | T2Agent: A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree SearchabstractReal-world multimodal misinformation often arises from mixed forgery sources, requiring dynamic reasoning and adaptive verification. However, existing methods mainly rely on static pipelines and limited tool usage, limiting their ability to handle such complexity and diversity. To address this challenge, we propose T2Agent, a novel misinformation detection agent that incorporates an extensible toolkit with Monte Carlo Tree Search (MCTS). The toolkit consists of modular tools such as web search, forgery detection, and consistency analysis. Each tool is described using standardized templates, enabling seamless integration and future expansion. To avoid inefficiency from using all tools simultaneously, a greedy search-based selector is proposed to identify a task-relevant subset. This subset then serves as the action space for MCTS to dynamically collect evidence and perform multi-source verification. To better align MCTS with the multi-source nature of misinformation detection, T2Agent extends traditional MCTS with multi-source verification, which decomposes the task into coordinated subtasks targeting different forgery sources. A dual reward mechanism containing a reasoning trajectory score and a confidence score is further proposed to encourage a balance between exploration across mixed forgery sources and exploitation for more reliable evidence. We conduct ablation studies to confirm the effectiveness of the tree search mechanism and tool usage. Extensive experiments further show that T2Agent consistently outperforms existing baselines on challenging mixed-source multimodal misinformation benchmarks, demonstrating its strong potential as a training-free detector. Xing Cui, Yueying Zou, Zekun Li 0001, Peipei Li 0002, Xuannan Liu, Huaibo Huang |
AAAI | 7 |
| 2026 | Uncertainty-Aware Source-Free Adaptive Image Restoration with State Space Augmentation
Yuang Ai, Jie Cao 0002, Ran He 0001, Huaibo Huang |
Int. J. Comput. Vis. | 4 |
| 2026 | Advancing Vision Transformer With Enhanced Spatial PriorsabstractIn recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and suffers from quadratic computational complexity, limiting its applicability. To address these issues, we have proposed RMT, a robust vision backbone with explicit spatial priors for general purposes. RMT utilizes Manhattan distance decay to introduce spatial information and employs a horizontal and vertical decomposition attention method to model global information. Building on the strengths of RMT, Euclidean enhanced Vision Transformer (EVT) is an expanded version that incorporates several key improvements. Firstly, EVT uses a more reasonable Euclidean distance decay to enhance the modeling of spatial information, allowing for a more accurate representation of spatial relationships compared to the Manhattan distance used in RMT. Secondly, EVT abandons the decomposed attention mechanism featured in RMT and instead adopts a simpler spatially-independent grouping approach, providing the model with greater flexibility in controlling the number of tokens within each group. By addressing these modifications, EVT offers a more sophisticated and adaptable approach to incorporating spatial priors into the Self-Attention mechanism, thus overcoming some of the limitations associated with RMT and further enhancing its applicability in various computer vision tasks. Extensive experiments on Image Classification, Object Detection, Instance Segmentation, and Semantic Segmentation demonstrate that EVT exhibits exceptional performance. Without additional training data, EVT achieves 86.6% top1-acc on ImageNet-1 k. Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Hongmin Liu 0001, Ran He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | InfoBFR: Real-World Blind Face Restoration via Information BottleneckabstractPrevious blind face restoration (BFR) methods have primarily leveraged facial priors from pretrained GAN or diffusion models. These neural BFR models suffer from diverse neural degradations, such as prior bias, topological distortion, textural distortion, and artifact residues, which limit their real-world generalization in complex real-world scenarios. In this paper, we propose an effective framework,InfoBFR, to address neural degradation from an information-theoretic perspective, which achieves BFR boosting in diverse wild and heterogeneous scenes. Specifically, on the basis of the results from pretrained BFR models, InfoBFR considers information compression by using a manifold information bottleneck (MIB) and manifold information compensation (MIC) with efficient diffusion LoRA to conduct information optimization. InfoBFR effectively synthesizes high-fidelity faces with texture and structure boosting. Comprehensive experimental results demonstrate the high boosting performance of InfoBFR (nearly 82%) for state-of-the-art GAN-based and diffusion-based BFR methods, as it can complete BFR tasks in approximately 70 ms and has 4M trainable parameters. It is promising that InfoBFR is the first unified postprocessing restorer universally employed by diverse BFR models to overcome the limitations of neural degradation. Nan Gao 0001, Jia Li 0044, Huaibo Huang, Ran He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Breaking the Low-Rank Dilemma of Linear AttentionabstractThe Softmax attention mechanism in Transformer models is notoriously computationally expensive due to its quadratic complexity, posing significant challenges in vision applications. In contrast, linear attention offers a far more efficient solution by reducing the complexity to linear levels. However, linear attention often suffers significant performance degradation compared to Softmax attention. Our experiments indicate that this performance drop stems from the low-rank nature of linear attention’s output feature map, which hinders its ability to adequately model complex spatial information. To address this low-rank dilemma, we conduct rank analysis from two perspectives: the KV buffer and the output features. Consequently, we introduce Rank-Augmented Linear Attention (RALA), which rivals the performance of Softmax attention while maintaining linear complexity and high efficiency. Building upon RALA, we construct the Rank-Augmented Vision Linear Transformer (RAVLT). Extensive experiments demonstrate that RAVLT achieves excellent performance across various vision tasks. Specifically, without using any additional labels, data, or supervision during training, RAVLT achieves an 84.4% Top-1 accuracy on ImageNet-1k with only 26M parameters and 4.6G FLOPs. This result significantly surpasses previous linear attention mechanisms, fully illustrating the potential of RALA. Code will be available at https://github.com/qhfan/RALA. Qihang Fan, Huaibo Huang, Ran He 0001 |
CVPR | 2 |
| 2025 | DM-DPR: Diffusion and Mamba-based Degradation Prediction for Blind Face RestorationabstractBlind Face Restoration (BFR), which involves converting low-quality facial images with unknown and varied degradation into high-quality counterparts, suffers from issues of sub-optimal restoration and over-correction due to inconsistent degradation levels. To rectify the above issue, inspired by the State Space model, especially the improved version Mamba’s enhanced long-range dependencies modeling ability and Stable Diffusion’s ability in integrating multi-modal prompts, we introduce a novel approach, Diff-Mamba Degradation Prediction Restoration (DM-DPR), to leverage the combination of a Mamba prompt generation framework with Stable Diffusion-based image restoration. Its core lies in two primary components: a Mamba-based multi-modal prompt generator that quantifies the degradation severity, generating corresponding textual and visual prompts, additionally with a multi-modal prompt driven Stable Diffusion process that adjusts restoration efforts based on the estimated degradation level. Derived from the CelebA-Test, we create degraded datasets exhibiting a wide range of degradation severity. Extensive experimental evaluations demonstrate that DM-DPR substantially surpasses existing state-of-the-art methods, thereby robustly establishing its enhanced capability to manage varying degrees of image degradation. Guorong Yuan, Huaibo Huang, Jie Cao 0002, Yuang Ai, Ran He 0001 |
FG | 3 |
| 2025 | Rectifying Magnitude Neglect in Linear AttentionabstractAs the core operator of Transformers, Softmax Attention exhibits excellent global modeling capabilities. However, its quadratic complexity limits its applicability to vision tasks. In contrast, Linear Attention shares a similar formulation with Softmax Attention while achieving linear complexity, enabling efficient global information modeling. Nevertheless, Linear Attention suffers from a significant performance degradation compared to standard Softmax Attention. In this paper, we analyze the underlying causes of this issue based on the formulation of Linear Attention. We find that, unlike Softmax Attention, Linear Attention entirely disregards the magnitude information of the Query. This prevents the attention score distribution from dynamically adapting as the Query scales. As a result, despite its structural similarity to Softmax Attention, Linear Attention exhibits a significantly different attention score distribution. Based on this observation, we propose Magnitude-Aware Linear Attention (MALA), which modifies the computation of Linear Attention to fully incorporate the Query's magnitude. This adjustment allows MALA to generate an attention score distribution that closely resembles Softmax Attention while exhibiting a more well-balanced structure. We evaluate the effectiveness of MALA on multiple tasks, including image classification, object detection, instance segmentation, semantic segmentation, natural language processing, speech recognition, and image generation. Our MALA achieves strong results on all of these tasks. Code will be available at https://github.com/qhfan/MALA Qihang Fan, Huaibo Huang, Yuang Ai, Ran He 0001 |
ICCV | 2 |
| 2025 | Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Ran He 0001 |
ICCV | 2 |
| 2025 | MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMsabstractCurrent multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this field. To address this, we introduce MMFakeBench, the first comprehensive benchmark for mixed-source MMD. MMFakeBench includes 3 critical sources: textual veracity distortion, visual veracity distortion, and cross-modal consistency distortion, along with 12 sub-categories of misinformation forgery types. We further conduct an extensive evaluation of 6 prevalent detection methods and 15 Large Vision-Language Models (LVLMs) on MMFakeBench under a zero-shot setting. The results indicate that current methods struggle under this challenging and realistic mixed-source MMD setting. Additionally, we propose MMD-Agent, a novel approach to integrate the reasoning, action, and tool-use capabilities of LVLM agents, significantly enhancing accuracy and generalization. We believe this study will catalyze future research into more realistic mixed-source multimodal misinformation and provide a fair evaluation of misinformation detection methods. Xuannan Liu, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Shuhan Xia, Xing Cui, Linzhi Huang, Weihong Deng, Zhaofeng He 0001 |
ICLR | 4 |
| 2025 | Degradation-Aware Multi-Task Image Restoration with State Space ModelsabstractImage restoration (IR) has made significant strides, evolving from basic pixel-wise restoration to more advanced techniques capable of handling diverse degradations. However, current all-in-one IR methods face challenges in effectively generalizing across complex real-world degradation scenarios. To address this challenge, we propose multi-task Restoration Mamba (ReMamba), a novel framework that leverages the power of state space models for long-sequence modeling to extract fine-grained degradation features from low-quality data. Guided by degradation-type prediction, which helps the model accurately identify and differentiate between various degradation types in the input, ReMamba employs prompt learning across both textual and visual modalities and enhances the restoration capabilities of downstream stable diffusion networks by injecting prompts as prior knowledge. Extensive experiments on both synthetic and real-world datasets, covering six distinct IR tasks, demonstrate the adaptability, generalizability, and robustness of ReMamba, highlighting its effectiveness in addressing a wide range of degradation types under challenging conditions. Purui Bai, Huaibo Huang, Jie Cao 0002, Yuang Ai, Ran He 0001 |
ICME | 3 |
| 2025 | Vision Transformer with Sparse Scan PriorabstractIn recent years, Transformers have achieved remarkable progress in computer vision tasks. However, their global modeling often comes with substantial computational overhead, in stark contrast to the human eye's efficient information processing. Inspired by the human eye's sparse scanning mechanism, we propose a Sparse Scan Self-Attention mechanism (S3A). This mechanism predefines a series of Anchors of Interest for each token and employs local attention to efficiently model the spatial information around these anchors, avoiding redundant global modeling and excessive focus on local information. This approach mirrors the human eye's functionality and significantly reduces the computational load of vision models. Building on S3A, we introduce the Sparse Scan Vision Transformer (SSViT). Extensive experiments demonstrate the outstanding performance of SSViT across a variety of tasks. Specifically, on ImageNet classification, without additional supervision or training data, SSViT achieves top-1 accuracies of 84.4%/85.7% with 4.4G/18.2G FLOPs. SSViT also excels in downstream tasks such as object detection, instance segmentation, and semantic segmentation. Its robustness is further validated across diverse datasets. Yuguang Zhang, Qihang Fan, Huaibo Huang |
ACM Multimedia | 3 |
| 2025 | DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion ModelingabstractDiffusion Transformer (DiT), a promising diffusion model for visual generation, demonstrates impressive performance but incurs significant computational overhead. Intriguingly, analysis of pre-trained DiT models reveals that global self-attention is often redundant, predominantly capturing local patterns—highlighting the potential for more efficient alternatives. In this paper, we revisit convolution as an alternative building block for constructing efficient and expressive diffusion models. However, naively replacing self-attention with convolution typically results in degraded performance. Our investigations attribute this performance gap to the higher channel redundancy in ConvNets compared to Transformers. To resolve this, we introduce a compact channel attention mechanism that promotes the activation of more diverse channels, thereby enhancing feature diversity. This leads to Diffusion ConvNet (DiCo), a family of diffusion models built entirely from standard ConvNet modules, offering strong generative performance with significant efficiency gains. On class-conditional ImageNet generation benchmarks, DiCo-XL achieves an FID of 2.05 at 256$\times$256 resolution and 2.53 at 512$\times$512, with a **2.7$\times$** and **3.1$\times$** speedup over DiT-XL/2, respectively. Furthermore, experimental results on MS-COCO demonstrate that the purely convolutional DiCo exhibits strong potential for text-to-image generation. Yuang Ai, Qihang Fan, Xuefeng Hu, Zhenheng Yang, Ran He 0001, Huaibo Huang |
NeurIPS | 6 |
| 2025 | Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMsabstractThe increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce distinct safety risks. To bridge this gap, we introduce Video-SafetyBench, the first comprehensive benchmark designed to evaluate the safety of LVLMs under video-text attacks. It comprises 2,264 video-text pairs spanning 48 fine-grained unsafe categories, each pairing a synthesized video with either a harmful query, which contains explicit malice, or a benign query, which appears harmless but triggers harmful behavior when interpreted alongside the video. To generate semantically accurate videos for safety evaluation, we design a controllable pipeline that decomposes video semantics into subject images (what is shown) and motion text (how it moves), which jointly guide the synthesis of query-relevant videos. To effectively evaluate uncertain or borderline harmful outputs, we propose RJScore, a novel LLM-based metric that incorporates the confidence of judge models and human-aligned decision threshold calibration. Extensive experiments show that benign-query video composition achieves average attack success rates of 67.2%, revealing consistent vulnerabilities to video-induced attacks. We believe Video-SafetyBench will catalyze future research into video-based safety evaluation and defense strategies. Xuannan Liu, Zekun Li 0001, Zheqi He, Peipei Li 0002, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang 0023, Ran He 0001 |
NeurIPS | 7 |
| 2025 | Straighter Flow Matching via a Diffusion-Based Coupling Prior
Siyu Xing, Jie Cao 0002, Huaibo Huang, Haichao Shi, Xiaoyu Zhang 0002 |
PRCV (8) | 3 |
| 2025 | Test-time Forgery Detection with Spatial-Frequency Prompt Learning
Junxian Duan, Yuang Ai, Shenyuan Huang, Huaibo Huang, Jie Cao 0002, Ran He 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | RealDTT: Towards A Comprehensive Real-World Dataset for Tampered Text Detection
Junxian Duan, Fan Ji, Zhiyong Wang 0001, Huaibo Huang |
Int. J. Comput. Vis. | 6 |
| 2025 | Select & Re-Rank: Effectively and efficiently matching multimodal data with dynamically evolving attention
Weikuo Guo, Xiangwei Kong 0001, Huaibo Huang |
Neurocomputing | 3 |
| 2025 | Deep learning technology for face forgery detection: A survey
Lixia Ma, Puning Yang, Zimin (Max) Yang, Huaibo Huang |
Neurocomputing | 6 |
| 2024 | Heterogeneous Test-Time Training for Multi-Modal Person Re-identificationabstractMulti-modal person re-identification (ReID) seeks to mitigate challenging lighting conditions by incorporating diverse modalities. Most existing multi-modal ReID methods concentrate on leveraging complementary multi-modal information via fusion or interaction. However, the relationships among heterogeneous modalities and the domain traits of unlabeled test data are rarely explored. In this paper, we propose a Heterogeneous Test-time Training (HTT) framework for multi-modal person ReID. We first propose a Cross-identity Inter-modal Margin (CIM) loss to amplify the differentiation among distinct identity samples. Moreover, we design a Multi-modal Test-time Training (MTT) strategy to enhance the generalization of the model by leveraging the relationships in the heterogeneous modalities and the information existing in the test data. Specifically, in the training stage, we utilize the CIM loss to further enlarge the distance between anchor and negative by forcing the inter-modal distance to maintain the margin, resulting in an enhancement of the discriminative capacity of the ultimate descriptor. Subsequently, since the test data contains characteristics of the target domain, we adapt the MTT strategy to optimize the network before the inference by using self-supervised tasks designed based on relationships among modalities. Experimental results on benchmark multi-modal ReID datasets RGBNT201, Market1501-MM, RGBN300, and RGBNT100 validate the effectiveness of the proposed method. The codes can be found at https://github.com/ziwang1121/HTT. Zi Wang 0013, Huaibo Huang, Aihua Zheng, Ran He 0001 |
AAAI | 2 |
| 2024 | DeVAn: Dense Video Annotation for Video-Language ModelsabstractTingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fang, Ding Zhou, Huaibo Huang, Ran He, Hongxia Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan, Huaibo Huang, Ran He 0001, Hongxia Yang |
ACL (1) | 6 |
| 2024 | Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image RestorationabstractDespite substantial progress, all-in-one image restoration (IR) grapples with persistent challenges in handling intricate real-world degradations. This paper introduces MPerceiver: a novel multimodal prompt learning approach that harnesses Stable Diffusion (SD) priors to enhance adaptiveness, generalizability and fidelity for all-in-one im-age restoration. Specifically, we develop a dual-branch module to master two types of SD prompts: textual for holistic representation and visual for multiscale detail rep-resentation. Both prompts are dynamically adjusted by degradation predictions from the CLIP image encoder, en-abling adaptive responses to diverse unknown degradations. Moreover, a plug-in detail refinement module im-proves restoration fidelity via direct encoder-to-decoder in-formation transformation. To assess our method, MPer-ceiver is trained on 9 tasks for all-in-one IR and outper-forms state-of-the-art task-specific methods across many tasks. Post multitask pre-training, MPerceiver attains a generalized representation in low-level vision, exhibiting remarkable zero-shot and few-shot capabilities in unseen tasks. Extensive experiments on 16 IR tasks underscore the superiority of MPerceiver in terms of adaptiveness, gener-alizability and fidelity. Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, Ran He 0001 |
CVPR | 2 |
| 2024 | Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation TransformerabstractUnsupervised Domain Adaptation (UDA) can effectively address domain gap issues in real-world image Super-Resolution (SR) by accessing both the source and target data. Considering privacy policies or transmission restrictions of source data in practical scenarios, we propose a SOurce-free Domain Adaptation framework for image SR (SODA-SR) to address this issue, i.e., adapt a source-trained model to a target domain with only unlabeled target data. SODA-SR leverages the source-trained model to generate refined pseudo-labels for teacher-student learning. To better utilize pseudo-labels, we propose a novel wavelet-based augmentation method, named Wavelet Augmentation Transformer (WAT), which can be flexibly incorporated with existing networks, to implicitly produce useful augmented data. WAT learns low-frequency information of varying levels across diverse samples, which is aggregated efficiently via deformable attention. Furthermore, an uncertainty-aware self-training mechanism is proposed to improve the accuracy of pseudo-labels, with inaccurate predictions being rectified by uncertainty estimation. To acquire better SR results and avoid overfitting pseudo-labels, several regularization losses are proposed to constrain target LR and SR images in the frequency domain. Experiments show that without accessing source data, SODA-SR outperforms state-of-the-art UDA methods in both synthetic→real and real→real adaptation settings, and is not constrained by specific network architectures. Yuang Ai, Xiaoqiang Zhou, Huaibo Huang, Ran He 0001 |
CVPR | 3 |
| 2024 | RMT: Retentive Networks Meet Vision TransformersabstractVision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. How-ever, the core component of ViT, Self-Attention, lacks ex-plicit spatial priors and bears a quadratic computational complexity, thereby constraining the applicability of ViT. To alleviate these issues, we draw inspiration from the re-cent Retentive Network (RetNet) in the field of NLP, and propose RMT, a strong vision backbone with explicit spa-tial prior for general purposes. Specifically, we extend the RetNet's temporal decay mechanism to the spatial do-main, and propose a spatial decay matrix based on the Manhattan distance to introduce the explicit spatial prior to Self-Attention. Additionally, an attention decomposition form that adeptly adapts to explicit spatial prior is proposed, aiming to reduce the computational burden of modeling global information without disrupting the spa-tial decay matrix. Based on the spatial decay matrix and the attention decomposition form, we can flexibly integrate explicit spatial prior into the vision backbone with lin-ear complexity. Extensive experiments demonstrate that RMT exhibits exceptional performance across various vision tasks. Specifically, without extra training data, RMT achieves 84.8% and 86.1% top-l acc on ImageNet-lk with 27MI4.5GFLOPs and 96M/18.2GFLOPs. For downstream tasks, RMT achieves 54.5 box AP and 47.2 mask AP on the COCO detection task, and 52.8 mloU on the ADE20K se-mantic segmentation task. Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Hongmin Liu 0001, Ran He 0001 |
CVPR | 2 |
| 2024 | INSTASTYLE: Inversion Noise of a Stylized Image is Secretly a Style Adviser
Xing Cui, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Xuannan Liu, Zhaofeng He 0001 |
ECCV (51) | 4 |
| 2024 | Semantic-Aware Detail Enhancement for Blind Face RestorationabstractThe goal of Blind Face Restoration is to recover high-quality images from low-quality images suffering from unknown degradations, posing a significantly challenging problem. In recent years, numerous BFR methods have been proposed, achieving significant success. However, faces possess a unique facial topology, and subtle differences in texture, slight structural imbalances, and minimal asymmetry are easily perceptible in the restored face images. Previous methods often struggle to generate realistically high-quality images from real-world low-quality images and fail to preserve fine features. To more effectively restore image details and textures, providing a more natural and realistic restoration effect, we integrate facial semantic information as prior knowledge into the blind face restoration task. We employ a multi-head cross-attention mechanism to simultaneously consider facial semantic information and context information for modeling. Additionally, we introduce a local detail enhancement module specifically designed to enhance the processing capability of details around the eyes and mouth. Experimental results indicate that our proposed method recovers facial images on synthetic and real datasets more realistically and with higher fidelity. Xiaoqiang Zhou, Jie Cao 0002, Huaibo Huang, Aihua Zheng, Ran He 0001 |
FG | 4 |
| 2024 | Parallel Augmentation and Dual Enhancement for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID), the task of searching for the same person’s images in occluded environments, has attracted lots of attention in the past decades. Recent approaches concentrate on improving performance on occluded data by data/feature augmentation or using extra models to predict occlusions. However, they ignore the imbalance problem in this task and can not fully utilize the information from the training data. To alleviate these two issues, we propose a simple yet effective method with Parallel Augmentation and Dual Enhancement (PADE), which is robust on both occluded and non-occluded data and does not require any auxiliary clues. First, we design a parallel augmentation mechanism (PAM) to generate more suitable occluded data to mitigate the negative effects of unbalanced data. Second, we propose the global and local dual enhancement strategy (DES) to promote the context information and details. Experimental results on three widely used occluded datasets and two non-occluded datasets validate the effectiveness of our method. The code is available at PADE (GitHub). Zi Wang 0013, Huaibo Huang, Aihua Zheng, Chenglong Li 0002, Ran He 0001 |
ICASSP | 2 |
| 2024 | ZePo: Zero-Shot Portrait Stylization with Faster SamplingabstractDiffusion-based text-to-image generation models have significantly advanced the field of art content synthesis. However, current portrait stylization methods generally require either model fine-tuning based on examples or the employment of DDIM Inversion to revert images to noise space, both of which substantially decelerate the image generation process. To overcome these limitations, this paper presents an inversion-free portrait stylization framework based on diffusion models that accomplishes content and style feature fusion in merely four sampling steps. We observed that Latent Consistency Models employing consistency distillation can effectively extract representative Consistency Features from noisy images. To blend the Consistency Features extracted from both content and style images, we introduce a Style Enhancement Attention Control technique that meticulously merges content and style features within the attention space of the target image. Moreover, we propose a feature merging strategy to amalgamate redundant features in Consistency Features, thereby reducing the computational load of attention control. Extensive experiments have validated the effectiveness of our proposed framework in enhancing stylization efficiency and fidelity. The code is available at \url{https://github.com/liujin112/ZePo}. Jin Liu 0040, Huaibo Huang, Jie Cao 0002, Ran He 0001 |
ACM Multimedia | 2 |
| 2024 | FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMsabstractThe massive generation of multimodal fake news involving both text and images exhibits substantial distribution discrepancies, prompting the need for generalized detectors. However, the insulated nature of training restricts the capability of classical detectors to obtain open-world facts. While Large Vision-Language Models (LVLMs) have encoded rich world knowledge, they are not inherently tailored for combating fake news and struggle to comprehend local forgery details. In this paper, we propose FKA-Owl, a novel framework that leverages forgery-specific knowledge to augment LVLMs, enabling them to reason about manipulations effectively. The augmented forgery-specific knowledge includes semantic correlation between text and images, and artifact trace in image manipulation. To inject these two kinds of knowledge into the LVLM, we design two specialized modules to establish their representations, respectively. The encoded knowledge embeddings are then incorporated into LVLMs. Extensive experiments on the public benchmark demonstrate that FKA-Owl achieves superior cross-domain performance compared to previous methods. Code is publicly available at https://liuxuannan.github.io/FKA_Owl.github.io/. Xuannan Liu, Peipei Li 0002, Huaibo Huang, Zekun Li 0001, Xing Cui, Lixiong Qin, Weihong Deng, Zhaofeng He 0001 |
ACM Multimedia | 3 |
| 2024 | DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset CurationabstractImage restoration (IR) in real-world scenarios presents significant challenges due to the lack of high-capacity models and comprehensive datasets.
To tackle these issues, we present a dual strategy: GenIR, an innovative data curation pipeline, and DreamClear, a cutting-edge Diffusion Transformer (DiT)-based image restoration model.
**GenIR**, our pioneering contribution, is a dual-prompt learning pipeline that overcomes the limitations of existing datasets, which typically comprise only a few thousand images and thus offer limited generalizability for larger models.
GenIR streamlines the process into three stages: image-text pair construction, dual-prompt based fine-tuning, and data generation \& filtering. This approach circumvents the laborious data crawling process, ensuring copyright compliance and providing a cost-effective, privacy-safe solution for IR dataset construction. The result is a large-scale dataset of one million high-quality images.
Our second contribution, **DreamClear**, is a DiT-based image restoration model. It utilizes the generative priors of text-to-image (T2I) diffusion models and the robust perceptual capabilities of multi-modal large language models (MLLMs) to achieve photorealistic restoration. To boost the model's adaptability to diverse real-world degradations, we introduce the Mixture of Adaptive Modulator (MoAM). It employs token-wise degradation priors to dynamically integrate various restoration experts, thereby expanding the range of degradations the model can address.
Our exhaustive experiments confirm DreamClear's superior performance, underlining the efficacy of our dual strategy for real-world image restoration. Code and pre-trained models are available at: https://github.com/shallowdream204/DreamClear. Yuang Ai, Xiaoqiang Zhou, Huaibo Huang, Quanzeng You, Hongxia Yang |
NeurIPS | 3 |
| 2024 | Visual Anchors Are Strong Information Aggregators For Multimodal Large Language ModelabstractIn the realm of Multimodal Large Language Models (MLLMs), vision-language connector plays a crucial role to link the pre-trained vision encoders with Large Language Models (LLMs). Despite its importance, the vision-language connector has been relatively less explored. In this study, we aim to propose a strong vision-language connector that enables MLLM to simultaneously achieve high accuracy and low computation cost. We first reveal the existence of the visual anchors in Vision Transformer and propose a cost-effective search algorithm to progressively extract them. Building on these findings, we introduce the Anchor Former (AcFormer), a novel vision-language connector designed to leverage the rich prior knowledge obtained from these visual anchors during pretraining, guiding the aggregation of information.
Through extensive experimentation, we demonstrate that the proposed method significantly reduces computational costs by nearly two-thirds, while simultaneously outperforming baseline methods. This highlights the effectiveness and efficiency of AcFormer. Haogeng Liu, Quanzeng You, Yongfei Liu, Huaibo Huang, Ran He 0001, Hongxia Yang |
NeurIPS | 5 |
| 2024 | Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content GenerationabstractRecent advancements in 3D content generation have been significant, primarily due to the visual priors provided by pretrained diffusion models. However, large 2D visual models exhibit spatial perception hallucinations, leading to multi-view inconsistency in 3D content generated through Score Distillation Sampling (SDS). This phenomenon, characterized by overfitting to specific views, is referred to as the "Janus Problem". In this work, we investigate the hallucination issues of pretrained models and find that large multimodal models without geometric constraints possess the capability to infer geometric structures, which can be utilized to mitigate multi-view inconsistency. Building on this, we propose a novel tuning-free method. We represent the multimodal inconsistency query information to detect specific hallucinations in 3D content, using this as an enhanced prompt to re-consist the 2D renderings of 3D and jointly optimize the structure and appearance across different views. Our approach does not require 3D training data and can be implemented plug-and-play within existing frameworks. Extensive experiments demonstrate that our method significantly improves the consistency of 3D content generation and specifically mitigates hallucinations caused by pretrained large models, achieving state-of-the-art performance compared to other optimization methods. Jie Cao 0002, Jin Liu 0040, Xiaoqiang Zhou, Huaibo Huang, Ran He 0001 |
NeurIPS | 5 |
| 2024 | Learning Fine-Grained and Semantically Aware Mamba Representations for Tampered Text Detection in Images
Jie Cao 0002, Huaibo Huang |
PRCV (7) | 6 |
| 2024 | ViPro-BEV: Few-Shot Visual Prompting for Bird's-Eye-View Perception
Guorong Yuan, Huaibo Huang, Qihang Fan |
PRCV (10) | 2 |
| 2024 | Variational Capsules for Image Analysis and Synthesis
Yuguang Zhang, Huaibo Huang |
PRCV (9) | 2 |
| 2024 | Uncertainty-aware image inpainting with adaptive feedback networkabstractWhile most image inpainting methods perform well on small image defects, they still struggle to deliver satisfactory results on large holes due to insufficient image guidance. To address this challenge, this paper proposes an uncertainty-aware adaptive feedback network (U2AFN), which incorporates an adaptive feedback mechanism to refine inpainting regions progressively. U2AFN predicts both an uncertainty map and an inpainting result simultaneously. During each iteration, the adaptive integration feedback block utilizes inpainting pixels with low uncertainty to guide the subsequent learning iteration. This process leads to a gradual reduction in uncertainty and produces more reliable inpainting outcomes. Our approach is extensively evaluated and compared on multiple datasets, demonstrating its superior performance over existing methods. The code is available at: https://codeocean.com/capsule/1901983/tree. Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Yaohui Wang 0001, Cunjian Chen |
Expert Syst. Appl. | 3 |
| 2024 | Dynamic Graph Memory Bank for Video InpaintingabstractA major challenge of the video inpainting task is aggregating spatial and temporal information in the corrupted video effectively. In this paper, we propose a dynamic graph memory bank to settle this challenge. To model the long-range temporal dependency, a memory bank is built and updated dynamically with the input visual information flow. The relationships among the memory items are modeled through a graph-based message propagation. Benefiting from the dynamic graph memory bank, both contents and their relationships in the corrupted video are well exploited as the inpainting process going on. Besides, the spatial misalignment across different frames may degrade the quality of features in the dynamic graph memory bank. To alleviate this issue, we propose a motion-guided feature alignment module. The proposed module cooperates with the dynamic graph memory bank to improve the network’s information aggregation ability in spatial and temporal dimensions. Extensive experiments on the YouTube-VOS and DAVIS datasets demonstrate the superiority of our approach when compared with the state-of-the-arts. Xiaoqiang Zhou, Chaoyou Fu, Huaibo Huang, Ran He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | RISTRA: Recursive Image Super-Resolution Transformer With Relativistic AssessmentabstractMany recent image restoration methods use Transformer as the backbone network and redesign the Transformer blocks. Differently, we explore the parameter-sharing mechanism over Transformer blocks and propose a dynamic recursive process to address the image super-resolution task efficiently. We firstly present a Recursive Image Super-resolution Transformer (RIST). By sharing the weights across different blocks, a plain forward process through the whole Transformer network can be folded into recursive iterations through a Transformer block. Such a parameter-sharing based recursive process can not only reduce the model size greatly, but also enable restoring images progressively. Features in the recursive process are modeled as a sequence and propagated with a temporal attention network. Besides, by analyzing the prediction variation across different iterations in RIST, we design a dynamic recursive process that can allocate adaptive computation costs to different samples. Specifically, a quality assessment network estimates the restoration quality and terminates the recursive process dynamically. We propose a relativistic learning strategy to simplify the objective from absolute image quality assessment to relativistic quality comparison. The proposed Recursive Image Super-resolution Transformer with Relativistic Assessment (RISTRA) reduces the model size greatly with the parameter-sharing mechanism, and achieves an instance-wise dynamic restoration process as well. Extensive experiments on several image super-resolution benchmarks show the superiority of our approach over state-of-the-art counterparts Xiaoqiang Zhou, Huaibo Huang, Zilei Wang, Ran He 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Pluralistic Aging Diffusion AutoencoderabstractFace aging is an ill-posed problem because multiple plausible aging patterns may correspond to a given input. Most existing methods often produce one deterministic estimation. This paper proposes a novel CLIP-driven Pluralistic Aging Diffusion Autoencoder (PADA) to enhance the diversity of aging patterns. First, we employ diffusion models to generate diverse low-level aging details via a sequential denoising reverse process. Second, we present Probabilistic Aging Embedding (PAE) to capture diverse high-level aging patterns, which represents age information as probabilistic distributions in the common CLIP latent space. A text-guided KL-divergence loss is designed to guide this learning. Our method can achieve pluralistic face aging conditioned on open-world aging texts and arbitrary unseen face images. Qualitative and quantitative experiments demonstrate that our method can generate more diverse and high-quality plausible aging results. Peipei Li 0002, Rui Wang 0124, Huaibo Huang, Ran He 0001, Zhaofeng He 0001 |
ICCV | 3 |
| 2023 | Lightweight Vision Transformer with Bidirectional InteractionabstractRecent advancements in vision backbones have significantly improved their performance by simultaneously modeling images’ local and global contexts. However, the bidirectional interaction between these two contexts has not been well explored and exploited, which is important in the human visual system. This paper proposes a **F**ully **A**daptive **S**elf-**A**ttention (FASA) mechanism for vision transformer to model the local and global information as well as the bidirectional interaction between them in context-aware ways. Specifically, FASA employs self-modulated convolutions to adaptively extract local representation while utilizing self-attention in down-sampled space to extract global representation. Subsequently, it conducts a bidirectional adaptation process between local and global representation to model their interaction. In addition, we introduce a fine-grained downsampling strategy to enhance the down-sampled self-attention mechanism for finer-grained global perception capability. Based on FASA, we develop a family of lightweight vision backbones, **F**ully **A**daptive **T**ransformer (FAT) family. Extensive experiments on multiple vision tasks demonstrate that FAT achieves impressive performance. Notably, FAT accomplishes a **77.6%** accuracy on ImageNet-1K using only **4.5M** parameters and **0.7G** FLOPs, which surpasses the most advanced ConvNets and Transformers with similar model size and computational costs. Moreover, our model exhibits faster speed on modern GPU compared to other models. Qihang Fan, Huaibo Huang, Xiaoqiang Zhou, Ran He 0001 |
NeurIPS | 2 |
| 2023 | Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal ClassificationabstractWe present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspire us to leverage the rich ordinal priors in human language by converting the original task into a vision-language alignment task. Consequently, we propose L2RCLIP, which fully utilizes the language priors from two perspectives. First, we introduce a complementary prompt tuning technique called RankFormer, designed to enhance the ordering relation of original rank prompts. It employs token-level attention with residual-style prompt blending in the word embedding space. Second, to further incorporate language priors, we revisit the approximate bound optimization of vanilla cross-entropy loss and restructure it within the cross-modal embedding space. Consequently, we propose a cross-modal ordinal pairwise loss to refine the CLIP feature space, where texts and images maintain both semantic alignment and ordering alignment. Extensive experiments on three ordinal classification tasks, including facial age estimation, historical color image (HCI) classification, and aesthetic assessment demonstrate its promising performance. Rui Wang 0124, Peipei Li 0002, Huaibo Huang, Chunshui Cao, Ran He 0001, Zhaofeng He 0001 |
NeurIPS | 3 |
| 2023 | Diverse features discovery transformer for pedestrian attribute recognition
Aihua Zheng, Jiaxiang Wang 0001, Huaibo Huang, Ran He 0001, Amir Hussain 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Memory Uncertainty Learning for Real-World Single Image DerainingabstractSingle image deraining has witnessed dramatic improvements by training deep neural networks on large-scale synthetic data. However, due to the discrepancy between authentic and synthetic rain images, it is challenging to directly extend existing methods to real-world scenes. To address this issue, we propose a memory-uncertainty guided semi-supervised method to learn rain properties simultaneously from synthetic and real data. The key aspect is developing a stochastic memory network that is equipped with memory modules to record prototypical rain patterns. The memory modules are updated in a self-supervised way, allowing the network to comprehensively capture rainy styles without the need for clean labels. The memory items are read stochastically according to their similarities with rain representations, leading to diverse predictions and efficient uncertainty estimation. Furthermore, we present an uncertainty-aware self-training mechanism to transfer knowledge from supervised deraining to unsupervised cases. An additional target network is adopted to produce pseudo-labels for unlabeled data, of which the incorrect ones are rectified by uncertainty estimates. Finally, we construct a new large-scale image deraining dataset of 10.2 k real rain images, significantly improving the diversity of real rain scenes. Experiments show that our method achieves more appealing results for real-world rain removal than recent state-of-the-art methods. Huaibo Huang, Mandi Luo, Ran He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Contextual Measures for Iris RecognitionabstractThe iris patterns of the human contain a large amount of randomly distributed and irregularly shaped microstructures. These microstructures make the human iris informative biometric traits. To learn identity representation from them, this paper regards each iris region as a potential microstructure and proposes contextual measures (CM) to model the correlations between them. CM adopts two parallel branches to learn global and local contexts in iris image. The first one is the globally contextual measure branch. It measures the global context involving the relationships between all regions for feature aggregation and is robust to local occlusions. Besides, we improve its spatial perception considering the positional randomness of the microstructures. The other one is the locally contextual measure branch. This branch considers the role of local details in the phenotypic distinctiveness of iris patterns and learns a series of relationship atoms to capture contextual information from a local perspective. In addition, we develop the perturbation bottleneck to make sure that the two branches learn divergent contexts. It introduces perturbation to limit the information flow from input images to identity features, forcing CM to learn discriminative contextual information for iris recognition. Experimental results suggest that global and local contexts are two different clues critical for accurate iris recognition. The superior performance on four benchmark iris datasets demonstrates the effectiveness of the proposed approach in within-database and cross-database scenarios. Jianze Wei, Yunlong Wang 0003, Huaibo Huang, Ran He 0001, Zhenan Sun, Xingyu Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Confidence-Calibrated Face Image Forgery Detection with Contrastive Representation Distillation
Puning Yang, Huaibo Huang, Zhiyong Wang 0001, Aijing Yu, Ran He 0001 |
ACCV (4) | 2 |
| 2022 | Rethinking Image Cropping: Exploring Diverse Compositions from Global ViewsabstractExisting image cropping works mainly use anchor evaluation methods or coordinate regression methods. However, it is difficult for pre-defined anchors to cover good crops globally, and the regression methods ignore the cropping diversity. In this paper, we regard image cropping as a set prediction problem. A set of crops regressed from multiple learnable anchors is matched with the labeled good crops, and a classifier is trained using the matching results to select a valid subset from all the predictions. This new perspective equips our model with globality and diversity, mitigating the shortcomings but inherit the strengthens of previous methods. Despite the advantages, the set prediction method causes inconsistency between the validity labels and the crops. To deal with this problem, we propose to smooth the validity labels with two different methods. The first method that uses crop qualities as direct guidance is designed for the datasets with nearly dense quality labels. The second method based on the self distillation can be used in sparsely labeled datasets. Experimental results on the public datasets show the merits of our approach over state-of-the-art counterparts. Gengyun Jia, Huaibo Huang, Chaoyou Fu, Ran He 0001 |
CVPR | 2 |
| 2022 | Artistic Style Discovery with Independent ComponentsabstractStyle transfer has been well studied in recent years with excellent performance processed. While existing methods usually choose CNNs as the powerful tool to accomplish superb stylization, less attention was paid to the latent style space. Rare exploration of underlying dimensions results in the poor style controllability and the limited practical application. In this work, we rethink the internal meaning of style features, further proposing a novel unsupervised algorithm for style discovery and achieving personalized manip-ulation. In particular, we take a closer look into the mechanism of style transfer and obtain different artistic style components from the latent space consisting of different style features. Then fresh styles can be generated by linear combination according to various style components. Experimental results have shown that our approach is superb in 1) restylizing the original output with the diverse artistic styles discovered from the latent space while keeping the content unchanged, and 2) being generic and compatible for various style transfer methods. Our code is available in this page: https://github.com/Shelsin/ArtIns. Yi Li 0018, Huaibo Huang, Haiyan Fu, Wanwan Wang, Yanqing Guo |
CVPR | 3 |
| 2022 | Fine-Grained Cross-Modal Retrieval with Triple-Streamed Memory Fusion Transformer EncoderabstractRecently, the powerful attention mechanism has been wildly used to learn the fine-grained cross-modal correspondences. However, the trade-off between effectiveness and efficiency sometimes bothers existing attention-mechanism-based meth-ods. To address this deficiency, we propose a novel Triple-streamed architecture with a newly designed Memory fusion Transformer Encoder (Tri-MTE) for fine-grained cross-modal retrieval. Specifically, the whole model reserves the “late fusion” strategy thus ensuring efficiency. To strengthen the inter-modality interaction and improve the effectiveness, a memory fusion stream is designed and inserted between the modality streams to remember the modality-irrelevant infor-mation. Encoding such information to the modality represen-tation would significantly enhance the cross-modal retrieval performance. Finally, a bionic memory activation constrain-t is proposed to aid the learning procedure. Extensive ex-periments on two benchmark datasets show that the proposed method achieves promising results. Weikuo Guo, Huaibo Huang, Xiangwei Kong 0001, Ran He 0001 |
ICME | 2 |
| 2022 | Orthogonal Transformer: An Efficient Vision Transformer Backbone with Token OrthogonalizationabstractWe present a general vision transformer backbone, called as Orthogonal Transformer, in pursuit of both efficiency and effectiveness. A major challenge for vision transformer is that self-attention, as the key element in capturing long-range dependency, is very computationally expensive for dense prediction tasks (e.g., object detection). Coarse global self-attention and local self-attention are then designed to reduce the cost, but they suffer from either neglecting local correlations or hurting global modeling. We present an orthogonal self-attention mechanism to alleviate these issues. Specifically, self-attention is computed in the orthogonal space that is reversible to the spatial domain but has much lower resolution. The capabilities of learning global dependency and exploring local correlations are maintained because every orthogonal token in self-attention can attend to the entire visual tokens. Remarkably, orthogonality is realized by constructing an endogenously orthogonal matrix that is friendly to neural networks and can be optimized as arbitrary orthogonal matrices. We also introduce Positional MLP to incorporate position information for arbitrary input resolutions as well as enhance the capacity of MLP. Finally, we develop a hierarchical architecture for Orthogonal Transformer. Extensive experiments demonstrate its strong performance on a broad range of vision tasks, including image classification, object detection, instance segmentation and semantic segmentation. Huaibo Huang, Xiaoqiang Zhou, Ran He 0001 |
NeurIPS | 1 |
| 2022 | Style-Based Attentive Network for Real-World Face Hallucination
Mandi Luo, Xin Ma 0031, Huaibo Huang, Ran He 0001 |
PRCV (4) | 3 |
| 2022 | Prior-Guided Multi-scale Fusion Transformer for Face Attribute Recognition
Shaoheng Song, Huaibo Huang, Jiaxiang Wang 0001, Aihua Zheng, Ran He 0001 |
PRCV (1) | 2 |
| 2022 | DVG-Face: Dual Variational Generation for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) refers to matching cross-domain faces and plays a crucial role in public security. Nevertheless, HFR is confronted with challenges from large domain discrepancy and insufficient heterogeneous data. In this paper, we formulate HFR as a dual generation problem, and tackle it via a novel dual variational generation (DVG-Face) framework. Specifically, a dual variational generator is elaborately designed to learn the joint distribution of paired heterogeneous images. However, the small-scale paired heterogeneous training data may limit the identity diversity of sampling. In order to break through the limitation, we propose to integrate abundant identity information of large-scale visible data into the joint distribution. Furthermore, a pairwise identity preserving loss is imposed on the generated paired heterogeneous images to ensure their identity consistency. As a consequence, massive new diverse paired heterogeneous images with the same identity can be generated from noises. The identity consistency and identity diversity properties allow us to employ these generated images to train the HFR network via a contrastive learning mechanism, yielding both domain-invariant and discriminative embedding features. Concretely, the generated paired heterogeneous images are regarded as positive pairs, and the images obtained from different samplings are considered as negative pairs. Our method achieves superior performances over state-of-the-art methods on seven challenging databases belonging to five HFR tasks, including NIR-VIS, Sketch-Photo, Profile-Frontal Photo, Thermal-VIS, and ID-Camera. Chaoyou Fu, Xiang Wu 0001, Yibo Hu 0001, Huaibo Huang, Ran He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Contrastive attention network with dense field estimation for face completion
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Zhenhua Chai, Xiaolin Wei |
Pattern Recognit. | 3 |
| 2022 | Memory-Modulated Transformer Network for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) aims at matching face images across different domains. It is challenging due to the severe domain discrepancies and overfitting caused by small training datasets. Some researchers apply a “recognition via generation” strategy and propose to solve the problem by translating images from a given domain into the visual domain. However, in many HFR tasks such as near-infrablack HFR, there is no paiblack data, which makes it an unsupervised generation. Pose variations, background differences, and many other factors present challenges. Moreover, the generated results lack diversity since many previous works regard this image translation as a “one-to-one” generation task. Considering the information deficiency in the input images, we propose to formulate this image translation process as a “one-to-many” generation problem. Specifically, we introduce reference images to guide the generation process. We propose a memory module to explore the prototypical style patterns of the reference domain. After self-supervised updating, the memory items are attentively aggregated to represent the style information. Moreover, to subtly fuse the contents of input images with the style of reference images, we propose a novel style transformer module. Specifically, we crop the encoded input and reference feature maps into patches, and use the style transformer to establish long-range dependencies between the input and reference patches. Thus, the style of every input patch is transferblack based on those of the most relevant reference patches. Extensive experiments on multiple datasets for various HFR tasks, including NIR-VIS, thermal-VIS, sketch-photo, and gray-RGB, are conducted. The robustness and effectiveness of the proposed MMTN are demonstrated both quantitatively and qualitatively. Mandi Luo, Haoxue Wu, Huaibo Huang, Weizan He, Ran He 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Towards More Discriminative and Robust Iris Recognition by Learning Uncertain FactorsabstractThe uncontrollable acquisition process limits the performance of iris recognition. In the acquisition process, various inevitable factors, including eyes, devices, and environment, hinder the iris recognition system from learning a discriminative identity representation. This leads to severe performance degradation. In this paper, we explore uncertain acquisition factors and propose uncertainty embedding (UE) and uncertainty-guided curriculum learning (UGCL) to mitigate the influence of acquisition factors. UE represents an iris image using a probabilistic distribution rather than a deterministic point (binary template or feature vector) that is widely adopted in iris recognition methods. Specifically, UE learns identity and uncertainty features from the input image, and encodes them as two independent components of the distribution, mean and variance. Based on this representation, an input image can be regarded as an instantiated feature sampled from the UE, and we can also generate various virtual features through sampling. UGCL is constructed by imitating the progressive learning process of newborns. Particularly, it selects virtual features to train the model in an easy-to-hard order at different training stages according to their uncertainty. In addition, an instance-level enhancement method is developed by utilizing local and global statistics to mitigate the data uncertainty from image noise and acquisition conditions in the pixel-level space. The experimental results on six benchmark iris datasets verify the effectiveness and generalization ability of the proposed method on same-sensor and cross-sensor recognition. Jianze Wei, Huaibo Huang, Yunlong Wang 0003, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Information Bottleneck Disentanglement for Identity SwappingabstractImproving the performance of face forgery detectors often requires more identity-swapped images of higher-quality. One core objective of identity swapping is to generate identity-discriminative faces that are distinct from the target while identical to the source. To this end, properly disentangling identity and identity-irrelevant information is critical and remains a challenging endeavor. In this work, we propose a novel information disentangling and swapping network, called InfoSwap, to extract the most expressive information for identity representation from a pre-trained face recognition model. The key insight of our method is to formulate the learning of disentangled representations as optimizing an information bottleneck tradeoff, in terms of finding an optimal compression of the pretrained latent features. Moreover, a novel identity contrastive loss is proposed for further disentanglement by requiring a proper distance between the generated identity and the target. While the most prior works have focused on using various loss functions to implicitly guide the learning of representations, we demonstrate that our model can provide explicit supervision for learning disentangled representations, achieving impressive performance in generating more identity-discriminative swapped faces. Gege Gao, Huaibo Huang, Chaoyou Fu, Ran He 0001 |
CVPR | 2 |
| 2021 | Memory Oriented Transfer Learning for Semi-Supervised Image DerainingabstractDeep learning based methods have shown dramatic improvements in image rain removal by using large-scale paired data of synthetic datasets. However, due to the various appearances of real rain streaks that may differ from those in the synthetic training data, it is challenging to directly extend existing methods to the real-world scenes. To address this issue, we propose a memory-oriented semi-supervised (MOSS) method which enables the network to explore and exploit the properties of rain streaks from both synthetic and real data. The key aspect of our method is designing an encoder-decoder neural network that is augmented with a self-supervised memory module, where items in the memory record the prototypical patterns of rain degradations and are updated in a self-supervised way. Consequently, the rainy styles can be comprehensively de-rived from synthetic or real-world degraded images without the need for clean labels. Furthermore, we present a self-training mechanism that attempts to transfer deraining knowledge from supervised rain removal to unsupervised cases. An additional target network, which is updated with an exponential moving average of the online deraining network, is utilized to produce pseudo-labels for unlabeled rainy images. Meanwhile, the deraining network is optimized with supervised objectives on both synthetic paired data and pseudo-paired noisy data. Extensive experiments show that the proposed method achieves more appealing results not only on limited labeled data but also on unlabeled real-world images than recent state-of-the-art methods. Huaibo Huang, Aijing Yu, Ran He 0001 |
CVPR | 1 |
| 2021 | Visual-Semantic Transformer for Face Forgery DetectionabstractThis paper proposes a novel Visual-Semantic Transformer (VST) to detect face forgery based on semantic aware feature relations. In face images, intrinsic feature relations exist between different semantic parsing regions. We find that face forgery algorithms always change such relations. Therefore, we start the approach by extracting Contextual Feature Sequence (CFS) using a transformer encoder to make the best abnormal feature relation patterns. Meanwhile, images are segmented as soft face regions by a face parsing module. Then we merge the CFS and the soft face regions as Visual Semantic Sequences (VSS) representing features of semantic regions. The VSS is fed into the transformer decoder, in which the relations in the semantic region level are modeled. Our method achieved 99.58% accuracy on FF++(Raw) and 96.16% accuracy on Celeb-DF. Extensive experiments demonstrate that our framework outperforms or is comparable with state-of-the-art detection methods, especially towards unseen forgery methods. Gengyun Jia, Huaibo Huang, Junxian Duan, Ran He 0001 |
IJCB | 3 |
| 2021 | Selective Wavelet Attention Learning for Single Image Deraining
Huaibo Huang, Aijing Yu, Zhenhua Chai, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 1 |
| 2021 | LAMP-HQ: A Large-Scale Multi-pose High-Quality Database and Benchmark for NIR-VIS Face Recognition
Aijing Yu, Haoxue Wu, Huaibo Huang, Zhen Lei 0001, Ran He 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | Informative Sample Mining Network for Multi-domain Image-to-Image Translation
Jie Cao 0002, Huaibo Huang, Yi Li 0018, Ran He 0001, Zhenan Sun |
ECCV (19) | 2 |
| 2020 | Hierarchical Face Aging Through Disentangled Latent Characteristics
Peipei Li 0002, Huaibo Huang, Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun |
ECCV (3) | 2 |
| 2020 | Free-Form Image Inpainting via Contrastive Attention NetworkabstractMost deep learning based image inpainting approaches adopt autoencoder or its variants to fill missing regions in images. Encoders are usually utilized to learn powerful representational spaces, which are important for dealing with sophisticated learning tasks. Specifically, in image inpainting tasks, masks with any shapes can appear anywhere in images (i.e., free-form masks) which form complex patterns. It is difficult for encoders to capture such powerful representations under this complex situation. To tackle this problem, we propose a self-supervised Siamese inference network to improve the robustness and generalization. It can encode contextual semantics from full resolution images and obtain more discriminative representations. we further propose a multi-scale decoder with a novel dual attention fusion module (DAF), which can combine both the restored and known regions in a smooth way. This multi-scale architecture is benefit for decoding discriminative representations learned by encoders into images layer by layer. In this way, unknown regions will be filled naturally from outside to inside. Qualitative and quantitative experiments on multiple datasets, including facial and natural datasets (i.e., Celeb-HQ, Pairs Street View, Places2 and ImageNet), demonstrate that our proposed method outperforms state-of-the-art methods in generating high-quality inpainting results. Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Zhenhua Chai, Xiaolin Wei, Ran He 0001 |
ICPR | 3 |
| 2020 | Attentional Wavelet Network for Traditional Chinese Painting TransferabstractTraditional Chinese paintings pay more attention to `Gongbi' and `Xieyi' in artworks, which raises a challenging task to generate Chinese paintings from photos. `Xieyi' creates high-level conception for paintings, while `Gongbi' refers to portraying local details in paintings. This paper proposes an attentional wavelet network for photo to Chinese painting transferring. We first introduce wavelets to obtain high-level conception and local details in Chinese paintings via 2-D haar wavelet transform. Moreover, we design high-level transform stream and local enhancement stream to dispose high frequencies and low frequency respectively. Furthermore, we exploit self-attention mechanism to compatibly pick up high-level information which is used to remedy the missing details when reconstructing the Chinese painting. To advance our experiment, we set up a new dataset named P2ADataset, with diverse photos and Chinese paintings on famous mountains around China. Experimental results comparing with the state-of-the-art style transferring algorithms verify the effectiveness of the proposed method. We will release the codes and data to the public. Rui Wang 0124, Huaibo Huang, Aihua Zheng, Ran He 0001 |
ICPR | 2 |
| 2020 | Exemplar Guided Cross-Spectral Face Hallucination via Mutual Information DisentanglementabstractRecently, many Near infrared-visible (NIR-VIS) heterogeneous face recognition (HFR) methods have been proposed in the community. But it remains a challenging problem because of the sensing gap along with large pose variations. In this paper, we propose an Exemplar Guided Cross-Spectral Face Hallucination (EGCH) to reduce the domain discrepancy through disentangled representation learning. For each modality, EGCH contains a spectral encoder as well as a structure encoder to disentangle spectral and structure representation, respectively. It also contains a traditional generator that reconstructs the input from the above two representations, and a structure generator that predicts the facial parsing map from the structure representation. Besides, mutual information minimization and maximization are conducted to boost disentanglement and make representations adequately expressed. Then the translation is built on structure representations between two modalities. Provided with the transformed NIR structure representation and original VIS spectral representation, EGCH is capable to produce high-fidelity VIS images that preserve the topology structure of the input NIR while transfer the spectral information of an arbitrary VIS exemplar. Extensive experiments demonstrate that the proposed method achieves more promising results both qualitatively and quantitatively than the state-of-the-art NIR-VIS methods. Haoxue Wu, Huaibo Huang, Aijing Yu, Jie Cao 0002, Zhen Lei 0001, Ran He 0001 |
ICPR | 2 |
| 2020 | Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence LearningabstractTalking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on either disentangling the information in a single image or learning temporal information between frames. However, cross-modality coherence between audio and video information has not been well addressed during synthesis. In this paper, we propose a novel arbitrary talking face generation framework by discovering the audio-visual coherence via the proposed Asymmetric Mutual Information Estimator (AMIE). In addition, we propose a Dynamic Attention (DA) block by selectively focusing the lip area of the input image during the training stage, to further enhance lip synchronization. Experimental results on benchmark LRW dataset and GRID dataset transcend the state-of-the-art methods on prevalent metrics with robust high-resolution synthesizing on gender and pose variations. Huaibo Huang, Yi Li 0018, Aihua Zheng, Ran He 0001 |
IJCAI | 2 |
| 2020 | Disentangled Representation Learning of Makeup Portraits in the Wild
Yi Li 0018, Huaibo Huang, Jie Cao 0002, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 2 |
| 2020 | A Survey of Deep Facial Attribute Analysis
Xin Zheng 0008, Yanqing Guo, Huaibo Huang, Yi Li 0018, Ran He 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | BLAN: Bi-directional ladder attentive network for facial attribute prediction
Xin Zheng 0008, Huaibo Huang, Yanqing Guo, Bo Wang 0024, Ran He 0001 |
Pattern Recognit. | 2 |
| 2019 | Disentangled Variational Representation for Heterogeneous Face RecognitionabstractVisible (VIS) to near infrared (NIR) face matching is a challenging problem due to the significant domain discrepancy between the domains and a lack of sufficient data for training cross-modal matching algorithms. Existing approaches attempt to tackle this problem by either synthesizing visible faces from NIR faces, extracting domain-invariant features from these modalities, or projecting heterogeneous data onto a common latent space for cross-modal matching. In this paper, we take a different approach in which we make use of the Disentangled Variational Representation (DVR) for crossmodal matching. First, we model a face representation with an intrinsic identity information and its within-person variations. By exploring the disentangled latent variable space, a variational lower bound is employed to optimize the approximate posterior for NIR and VIS representations. Second, aiming at obtaining more compact and discriminative disentangled latent space, we impose a minimization of the identity information for the same subject and a relaxed correlation alignment constraint between the NIR and VIS modality variations. An alternative optimization scheme is proposed for the disentangled variational representation part and the heterogeneous face recognition network part. The mutual promotion between these two parts effectively reduces the NIR and VIS domain discrepancy and alleviates over-fitting. Extensive experiments on three challenging NIR-VIS heterogeneous face recognition databases demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods. Xiang Wu 0001, Huaibo Huang, Vishal M. Patel, Ran He 0001, Zhenan Sun |
AAAI | 2 |
| 2019 | Learning Disentangled Representation for Cross-Modal Retrieval with Deep Mutual Information EstimationabstractCross-modal retrieval has become a hot research topic in recent years for its theoretical and practical significance. This paper proposes a new technique for learning such deep visual-semantic embedding that is more effective and interpretable for cross-modal retrieval. The proposed method employs a two-stage strategy to fulfill the task. In the first stage, deep mutual information estimation is incorporated into the objective to maximize the mutual information between the input data and its embedding. In the second stage, an expelling branch is added to the network to disentangle the modality-exclusive information from the learned representations. This helps to reduce the impact of modality-exclusive information to the common subspace representation as well as improve the interpretability of the learned feature. Extensive experiments on two large-scale benchmark datasets demonstrate that our method can learn better visual-semantic embedding and achieve state-of-the-art cross-modal retrieval results. Weikuo Guo, Huaibo Huang, Xiangwei Kong 0001, Ran He 0001 |
ACM Multimedia | 2 |
| 2019 | Dual Variational Generation for Low Shot Heterogeneous Face RecognitionabstractHeterogeneous Face Recognition (HFR) is a challenging issue because of the large domain discrepancy and a lack of heterogeneous data. This paper considers HFR as a dual generation problem, and proposes a novel Dual Variational Generation (DVG) framework. It generates large-scale new paired heterogeneous images with the same identity from noise, for the sake of reducing the domain gap of HFR. Specifically, we first introduce a dual variational autoencoder to represent a joint distribution of paired heterogeneous images. Then, in order to ensure the identity consistency of the generated paired heterogeneous images, we impose a distribution alignment in the latent space and a pairwise identity preserving in the image space. Moreover, the HFR network reduces the domain discrepancy by constraining the pairwise feature distances between the generated paired heterogeneous images. Extensive experiments on four HFR databases show that our method can significantly improve state-of-the-art results. When using the generated paired images for training, our method gains more than 18\% True Positive Rate improvements over the baseline model when False Positive Rate is at $10^{-5}$. Chaoyou Fu, Xiang Wu 0001, Yibo Hu 0001, Huaibo Huang, Ran He 0001 |
NeurIPS | 4 |
| 2019 | Wavelet Domain Generative Adversarial Network for Multi-scale Face Hallucination
Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 1 |
| 2018 | IntroVAE: Introspective Variational Autoencoders for Photographic Image SynthesisabstractWe present a novel introspective variational autoencoder (IntroVAE) model for synthesizing high-resolution photographic images. IntroVAE is capable of self-evaluating the quality of its generated samples and improving itself accordingly. Its inference and generator models are jointly trained in an introspective way. On one hand, the generator is required to reconstruct the input images from the noisy outputs of the inference model as normal VAEs. On the other hand, the inference model is encouraged to classify between the generated and real samples while the generator tries to fool it as GANs. These two famous generative frameworks are integrated in a simple yet efficient single-stream architecture that can be trained in a single stage. IntroVAE preserves the advantages of VAEs, such as stable training and nice latent manifold. Unlike most other hybrid models of VAEs and GANs, IntroVAE requires no extra discriminators, because the inference model itself serves as a discriminator to distinguish between the generated and real samples. Experiments demonstrate that our method produces high-resolution photo-realistic images (e.g., CELEBA images at (1024^{2})), which are comparable to or better than the state-of-the-art GANs. Huaibo Huang, Zhihang Li, Ran He 0001, Zhenan Sun, Tieniu Tan |
NeurIPS | 1 |
| 2017 | Wavelet-SRNet: A Wavelet-Based CNN for Multi-scale Face Super ResolutionabstractMost modern face super-resolution methods resort to convolutional neural networks (CNN) to infer highresolution (HR) face images. When dealing with very low resolution (LR) images, the performance of these CNN based methods greatly degrades. Meanwhile, these methods tend to produce over-smoothed outputs and miss some textural details. To address these challenges, this paper presents a wavelet-based CNN approach that can ultra-resolve a very low resolution face image of 16 × 16 or smaller pixelsize to its larger version of multiple scaling factors (2×, 4×, 8× and even 16×) in a unified framework. Different from conventional CNN methods directly inferring HR images, our approach firstly learns to predict the LR's corresponding series of HR's wavelet coefficients before reconstructing HR images from them. To capture both global topology information and local texture details of human faces, we present a flexible and extensible convolutional neural network with three types of loss: wavelet prediction loss, texture loss and full-image loss. Extensive experiments demonstrate that the proposed approach achieves more appealing results both quantitatively and qualitatively than state-ofthe- art super-resolution methods. Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan |
ICCV | 1 |