Xiaoqiang Zhou

dblp:13/1515 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 UniAnimate: taming unified video diffusion models for consistent human image animation
Xiang Wang 0012, Shiwei Zhang 0001, Changxin Gao, Xiaoqiang Zhou, Yingya Zhang, Luxin Yan, Nong Sang
Sci. China Inf. Sci.5
2024 Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration
abstract
Despite substantial progress, all-in-one image restoration (IR) grapples with persistent challenges in handling intricate real-world degradations. This paper introduces MPerceiver: a novel multimodal prompt learning approach that harnesses Stable Diffusion (SD) priors to enhance adaptiveness, generalizability and fidelity for all-in-one im-age restoration. Specifically, we develop a dual-branch module to master two types of SD prompts: textual for holistic representation and visual for multiscale detail rep-resentation. Both prompts are dynamically adjusted by degradation predictions from the CLIP image encoder, en-abling adaptive responses to diverse unknown degradations. Moreover, a plug-in detail refinement module im-proves restoration fidelity via direct encoder-to-decoder in-formation transformation. To assess our method, MPer-ceiver is trained on 9 tasks for all-in-one IR and outper-forms state-of-the-art task-specific methods across many tasks. Post multitask pre-training, MPerceiver attains a generalized representation in low-level vision, exhibiting remarkable zero-shot and few-shot capabilities in unseen tasks. Extensive experiments on 16 IR tasks underscore the superiority of MPerceiver in terms of adaptiveness, gener-alizability and fidelity.
Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, Ran He 0001
CVPR3
2024 Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation Transformer
abstract
Unsupervised Domain Adaptation (UDA) can effectively address domain gap issues in real-world image Super-Resolution (SR) by accessing both the source and target data. Considering privacy policies or transmission restrictions of source data in practical scenarios, we propose a SOurce-free Domain Adaptation framework for image SR (SODA-SR) to address this issue, i.e., adapt a source-trained model to a target domain with only unlabeled target data. SODA-SR leverages the source-trained model to generate refined pseudo-labels for teacher-student learning. To better utilize pseudo-labels, we propose a novel wavelet-based augmentation method, named Wavelet Augmentation Transformer (WAT), which can be flexibly incorporated with existing networks, to implicitly produce useful augmented data. WAT learns low-frequency information of varying levels across diverse samples, which is aggregated efficiently via deformable attention. Furthermore, an uncertainty-aware self-training mechanism is proposed to improve the accuracy of pseudo-labels, with inaccurate predictions being rectified by uncertainty estimation. To acquire better SR results and avoid overfitting pseudo-labels, several regularization losses are proposed to constrain target LR and SR images in the frequency domain. Experiments show that without accessing source data, SODA-SR outperforms state-of-the-art UDA methods in both synthetic→real and real→real adaptation settings, and is not constrained by specific network architectures.
Yuang Ai, Xiaoqiang Zhou, Huaibo Huang, Ran He 0001
CVPR2
2024 Semantic-Aware Detail Enhancement for Blind Face Restoration
abstract
The goal of Blind Face Restoration is to recover high-quality images from low-quality images suffering from unknown degradations, posing a significantly challenging problem. In recent years, numerous BFR methods have been proposed, achieving significant success. However, faces possess a unique facial topology, and subtle differences in texture, slight structural imbalances, and minimal asymmetry are easily perceptible in the restored face images. Previous methods often struggle to generate realistically high-quality images from real-world low-quality images and fail to preserve fine features. To more effectively restore image details and textures, providing a more natural and realistic restoration effect, we integrate facial semantic information as prior knowledge into the blind face restoration task. We employ a multi-head cross-attention mechanism to simultaneously consider facial semantic information and context information for modeling. Additionally, we introduce a local detail enhancement module specifically designed to enhance the processing capability of details around the eyes and mouth. Experimental results indicate that our proposed method recovers facial images on synthetic and real datasets more realistically and with higher fidelity.
Xiaoqiang Zhou, Jie Cao 0002, Huaibo Huang, Aihua Zheng, Ran He 0001
FG2
2024 DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset Curation
abstract
Image restoration (IR) in real-world scenarios presents significant challenges due to the lack of high-capacity models and comprehensive datasets. To tackle these issues, we present a dual strategy: GenIR, an innovative data curation pipeline, and DreamClear, a cutting-edge Diffusion Transformer (DiT)-based image restoration model. **GenIR**, our pioneering contribution, is a dual-prompt learning pipeline that overcomes the limitations of existing datasets, which typically comprise only a few thousand images and thus offer limited generalizability for larger models. GenIR streamlines the process into three stages: image-text pair construction, dual-prompt based fine-tuning, and data generation \& filtering. This approach circumvents the laborious data crawling process, ensuring copyright compliance and providing a cost-effective, privacy-safe solution for IR dataset construction. The result is a large-scale dataset of one million high-quality images. Our second contribution, **DreamClear**, is a DiT-based image restoration model. It utilizes the generative priors of text-to-image (T2I) diffusion models and the robust perceptual capabilities of multi-modal large language models (MLLMs) to achieve photorealistic restoration. To boost the model's adaptability to diverse real-world degradations, we introduce the Mixture of Adaptive Modulator (MoAM). It employs token-wise degradation priors to dynamically integrate various restoration experts, thereby expanding the range of degradations the model can address. Our exhaustive experiments confirm DreamClear's superior performance, underlining the efficacy of our dual strategy for real-world image restoration. Code and pre-trained models are available at: https://github.com/shallowdream204/DreamClear.
Yuang Ai, Xiaoqiang Zhou, Huaibo Huang, Quanzeng You, Hongxia Yang
NeurIPS2
2024 Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation
abstract
Recent advancements in 3D content generation have been significant, primarily due to the visual priors provided by pretrained diffusion models. However, large 2D visual models exhibit spatial perception hallucinations, leading to multi-view inconsistency in 3D content generated through Score Distillation Sampling (SDS). This phenomenon, characterized by overfitting to specific views, is referred to as the "Janus Problem". In this work, we investigate the hallucination issues of pretrained models and find that large multimodal models without geometric constraints possess the capability to infer geometric structures, which can be utilized to mitigate multi-view inconsistency. Building on this, we propose a novel tuning-free method. We represent the multimodal inconsistency query information to detect specific hallucinations in 3D content, using this as an enhanced prompt to re-consist the 2D renderings of 3D and jointly optimize the structure and appearance across different views. Our approach does not require 3D training data and can be implemented plug-and-play within existing frameworks. Extensive experiments demonstrate that our method significantly improves the consistency of 3D content generation and specifically mitigates hallucinations caused by pretrained large models, achieving state-of-the-art performance compared to other optimization methods.
Jie Cao 0002, Jin Liu 0040, Xiaoqiang Zhou, Huaibo Huang, Ran He 0001
NeurIPS4
2024 TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection
Wenkui Yang, Xiaoqiang Zhou, Junxian Duan, Jie Cao 0002
PRCV (11)3
2024 Uncertainty-aware image inpainting with adaptive feedback network
abstract
While most image inpainting methods perform well on small image defects, they still struggle to deliver satisfactory results on large holes due to insufficient image guidance. To address this challenge, this paper proposes an uncertainty-aware adaptive feedback network (U2AFN), which incorporates an adaptive feedback mechanism to refine inpainting regions progressively. U2AFN predicts both an uncertainty map and an inpainting result simultaneously. During each iteration, the adaptive integration feedback block utilizes inpainting pixels with low uncertainty to guide the subsequent learning iteration. This process leads to a gradual reduction in uncertainty and produces more reliable inpainting outcomes. Our approach is extensively evaluated and compared on multiple datasets, demonstrating its superior performance over existing methods. The code is available at: https://codeocean.com/capsule/1901983/tree.
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Yaohui Wang 0001, Cunjian Chen
Expert Syst. Appl.2
2024 Dynamic Graph Memory Bank for Video Inpainting
abstract
A major challenge of the video inpainting task is aggregating spatial and temporal information in the corrupted video effectively. In this paper, we propose a dynamic graph memory bank to settle this challenge. To model the long-range temporal dependency, a memory bank is built and updated dynamically with the input visual information flow. The relationships among the memory items are modeled through a graph-based message propagation. Benefiting from the dynamic graph memory bank, both contents and their relationships in the corrupted video are well exploited as the inpainting process going on. Besides, the spatial misalignment across different frames may degrade the quality of features in the dynamic graph memory bank. To alleviate this issue, we propose a motion-guided feature alignment module. The proposed module cooperates with the dynamic graph memory bank to improve the network’s information aggregation ability in spatial and temporal dimensions. Extensive experiments on the YouTube-VOS and DAVIS datasets demonstrate the superiority of our approach when compared with the state-of-the-arts.
Xiaoqiang Zhou, Chaoyou Fu, Huaibo Huang, Ran He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 RISTRA: Recursive Image Super-Resolution Transformer With Relativistic Assessment
abstract
Many recent image restoration methods use Transformer as the backbone network and redesign the Transformer blocks. Differently, we explore the parameter-sharing mechanism over Transformer blocks and propose a dynamic recursive process to address the image super-resolution task efficiently. We firstly present a Recursive Image Super-resolution Transformer (RIST). By sharing the weights across different blocks, a plain forward process through the whole Transformer network can be folded into recursive iterations through a Transformer block. Such a parameter-sharing based recursive process can not only reduce the model size greatly, but also enable restoring images progressively. Features in the recursive process are modeled as a sequence and propagated with a temporal attention network. Besides, by analyzing the prediction variation across different iterations in RIST, we design a dynamic recursive process that can allocate adaptive computation costs to different samples. Specifically, a quality assessment network estimates the restoration quality and terminates the recursive process dynamically. We propose a relativistic learning strategy to simplify the objective from absolute image quality assessment to relativistic quality comparison. The proposed Recursive Image Super-resolution Transformer with Relativistic Assessment (RISTRA) reduces the model size greatly with the parameter-sharing mechanism, and achieves an instance-wise dynamic restoration process as well. Extensive experiments on several image super-resolution benchmarks show the superiority of our approach over state-of-the-art counterparts
Xiaoqiang Zhou, Huaibo Huang, Zilei Wang, Ran He 0001
IEEE Trans. Multim.1
2023 Lightweight Vision Transformer with Bidirectional Interaction
abstract
Recent advancements in vision backbones have significantly improved their performance by simultaneously modeling images’ local and global contexts. However, the bidirectional interaction between these two contexts has not been well explored and exploited, which is important in the human visual system. This paper proposes a **F**ully **A**daptive **S**elf-**A**ttention (FASA) mechanism for vision transformer to model the local and global information as well as the bidirectional interaction between them in context-aware ways. Specifically, FASA employs self-modulated convolutions to adaptively extract local representation while utilizing self-attention in down-sampled space to extract global representation. Subsequently, it conducts a bidirectional adaptation process between local and global representation to model their interaction. In addition, we introduce a fine-grained downsampling strategy to enhance the down-sampled self-attention mechanism for finer-grained global perception capability. Based on FASA, we develop a family of lightweight vision backbones, **F**ully **A**daptive **T**ransformer (FAT) family. Extensive experiments on multiple vision tasks demonstrate that FAT achieves impressive performance. Notably, FAT accomplishes a **77.6%** accuracy on ImageNet-1K using only **4.5M** parameters and **0.7G** FLOPs, which surpasses the most advanced ConvNets and Transformers with similar model size and computational costs. Moreover, our model exhibits faster speed on modern GPU compared to other models.
Qihang Fan, Huaibo Huang, Xiaoqiang Zhou, Ran He 0001
NeurIPS3
2023 Towards Lightweight Pixel-Wise Hallucination for Heterogeneous Face Recognition
abstract
Cross-spectral face hallucination is an intuitive way to mitigate the modality discrepancy in Heterogeneous Face Recognition (HFR). However, due to imaging differences, the hallucination inevitably suffers from a shape misalignment between paired heterogeneous images. Rather than building complicated architectures to circumvent the problem like previous works, we propose a simple yet effective method called Shape Alignment FacE (SAFE). Specifically, given an image, we align its shape to that of the paired one under the assistance of a 3D face model. The produced aligned pair enables us to train a lightweight generator that solely concentrates on spectrum translation with a pixel-wise supervision. However, since the 3D face model is powerless to attributes like the hair and glasses, there are still pixel discrepancies between the aligned pair. Given that, in the image space, we introduce a probabilistic pixel-wise loss that incorporates the discrepancies into a probabilistic distribution. Moreover, in order to alleviate the influence of the shape misalignment on spectrum translation, a spectrum optimal transport is performed in a shape-irrelevant latent space. Note that, in the final inference phase, except the lightweight generator, all other auxiliary modules are discarded. In addition to superior performance in qualitative synthesis and quantitative recognition, extensive experiments on 6 datasets demonstrate that our method also gains other two distinct advantages over existing state-of-the-art counterparts. The first is using a more lightweight generator. Compared with the state-of-the-art method, our method can achieve higher recognition results with 128x fewer parameters and 63x fewer FLOPs with only 4.58 ms latency on a single TITAN-XP. The second is training on low-shot datasets such as Oulu-CASIA NIR-VIS that just contains 1,920 images from 20 identities. To the best of our knowledge, we are the first that can perform well on such a small-scale dataset. These advantages make our method more practical in the real world and further push boundaries of heterogeneous face recognition.
Chaoyou Fu, Xiaoqiang Zhou, Weizan He, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Orthogonal Transformer: An Efficient Vision Transformer Backbone with Token Orthogonalization
abstract
We present a general vision transformer backbone, called as Orthogonal Transformer, in pursuit of both efficiency and effectiveness. A major challenge for vision transformer is that self-attention, as the key element in capturing long-range dependency, is very computationally expensive for dense prediction tasks (e.g., object detection). Coarse global self-attention and local self-attention are then designed to reduce the cost, but they suffer from either neglecting local correlations or hurting global modeling. We present an orthogonal self-attention mechanism to alleviate these issues. Specifically, self-attention is computed in the orthogonal space that is reversible to the spatial domain but has much lower resolution. The capabilities of learning global dependency and exploring local correlations are maintained because every orthogonal token in self-attention can attend to the entire visual tokens. Remarkably, orthogonality is realized by constructing an endogenously orthogonal matrix that is friendly to neural networks and can be optimized as arbitrary orthogonal matrices. We also introduce Positional MLP to incorporate position information for arbitrary input resolutions as well as enhance the capacity of MLP. Finally, we develop a hierarchical architecture for Orthogonal Transformer. Extensive experiments demonstrate its strong performance on a broad range of vision tasks, including image classification, object detection, instance segmentation and semantic segmentation.
Huaibo Huang, Xiaoqiang Zhou, Ran He 0001
NeurIPS2
2022 Contrastive attention network with dense field estimation for face completion
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Zhenhua Chai, Xiaolin Wei
Pattern Recognit.2
2020 Free-Form Image Inpainting via Contrastive Attention Network
abstract
Most deep learning based image inpainting approaches adopt autoencoder or its variants to fill missing regions in images. Encoders are usually utilized to learn powerful representational spaces, which are important for dealing with sophisticated learning tasks. Specifically, in image inpainting tasks, masks with any shapes can appear anywhere in images (i.e., free-form masks) which form complex patterns. It is difficult for encoders to capture such powerful representations under this complex situation. To tackle this problem, we propose a self-supervised Siamese inference network to improve the robustness and generalization. It can encode contextual semantics from full resolution images and obtain more discriminative representations. we further propose a multi-scale decoder with a novel dual attention fusion module (DAF), which can combine both the restored and known regions in a smooth way. This multi-scale architecture is benefit for decoding discriminative representations learned by encoders into images layer by layer. In this way, unknown regions will be filled naturally from outside to inside. Qualitative and quantitative experiments on multiple datasets, including facial and natural datasets (i.e., Celeb-HQ, Pairs Street View, Places2 and ImageNet), demonstrate that our proposed method outperforms state-of-the-art methods in generating high-quality inpainting results.
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Zhenhua Chai, Xiaolin Wei, Ran He 0001
ICPR2
2020 Image Inpainting with Contrastive Relation Network
abstract
Image inpainting faces the challenging issue of the requirements on structure reasonableness and texture coherence. In this paper, we propose a two-stage inpainting framework to address this issue. The basic idea is to address the two requirements in two separate stages. Completed segmentation of the corrupted image is firstly predicted through segmentation reconstruction network, while fine-grained image details are restored in the second stage through an image generator. The two stages are connected in series as the image details are generated under the guidance of completed segmentation map that predicted in the first stage. Specifically, in the second stage, we propose a novel graph-based relation network to model the relationship existed in corrupted image. In relation network, both intra-relationship for pixels in the same semantic region and inter-relationship between different semantic parts are considered, improving the consistency and compatibility of image textures. Besides, contrastive loss is designed to facilitate the relation network training. Such a framework not only simplifies the inpainting problem directly, but also exploits the relationship in corrupted image explicitly. Extensive experiments on various public datasets quantitatively and qualitatively demonstrate the superiority of our approach compared with the state-of-the-art.
Xiaoqiang Zhou, Junjie Li 0002, Zilei Wang, Ran He 0001, Tieniu Tan
ICPR1
2018 Recurrent convolutional neural network for answer selection in community question answering
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001
Neurocomputing1
2016 Incorporating Label Dependency for Answer Quality Tagging in Community Question Answering via CNN-LSTM-CRF
abstract
In community question answering (cQA), the quality of answers are determined by the matching degree between question-answer pairs and the correlation among the answers. In this paper, we show that the dependency between the answer quality labels also plays a pivotal role. To validate the effectiveness of label dependency, we propose two neural network-based models, with different combination modes of Convolutional Neural Net-works, Long Short Term Memory and Conditional Random Fields. Extensive experi-ments are taken on the dataset released by the SemEval-2015 cQA shared task. The first model is a stacked ensemble of the networks. It achieves 58.96% on macro averaged F1, which improves the state-of-the-art neural network-based method by 2.82% and outper-forms the Top-1 system in the shared task by 1.77%. The second is a simple attention-based model whose input is the connection of the question and its corresponding answers. It produces promising results with 58.29% on overall F1 and gains the best performance on the Good and Bad categories.
Yang Xiang 0003, Xiaoqiang Zhou, Qingcai Chen, Zhihui Zheng, Buzhou Tang, Xiaolong Wang 0001, Yang Qin 0001
COLING2
2016 An in-network caching scheme based on betweenness and content popularity prediction in content-centric networking
abstract
Content-centric Networking (CCN) is considered as a promising architecture to achieve reliable content distribution at large scale. One of the key research items of CCN is cache strategy, and most of the existing approaches consider little of the dynamicity of user interests. In this paper, we present a new cache policy, named as the betweenness and content popularity prediction (BEACON). Betweenness measures the importance of nodes in the whole network, and content popularity represents the user preference for service contents. By taking into account both network topology characteristics and flow distribution, the load of network and server is optimized. Moreover, we use the gray model to predict the content popularity, tracking the trend of user interest. The simulation results demonstrate that the BEACON scheme can effectively improve the cache hit rate, shorten access distance and reduce the delay of transmission.
Xiaoqiang Zhou, Min Zhao 0002, Muqing Wu
PIMRC1
2015 An Auto-Encoder for Learning Conversation Representation Using LSTM
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001
ICONIP (1)1
2013 Grammatical Error Correction Using Feature Selection and Confidence Tuning
Yang Xiang 0003, Yaoyun Zhang, Xiaolong Wang 0001, Chongqiang Wei, Xiaoqiang Zhou, Yuxiu Hu, Yang Qin 0001
IJCNLP6
2001 Authenticity and Integrity of Digital Mammography Images
abstract
Data security becomes more and more important in telemammography which uses a public high-speed wide area network connecting the examination site with the mammography expert center. Generally, security is characterized in terms of privacy, authenticity and integrity of digital data. Privacy is a network access issue and is not considered in this paper. We present a method, authenticity and integrity of digital mammography, here which can meet the requirements of authenticity and integrity for mammography image (IM) transmission. The authenticity and integrity for mammography (AIDM) consists of the following four modules. 1) Image preprocessing: To segment breast pixels from background and extract patient information from digital imaging and communication in medicine (DICOM) image header. 2) Image hashing: To compute an image hash value of the mammogram using the MD5 hash algorithm. 3) Data encryption: To produce a digital envelope containing the encrypted image hash value (digital signature) and corresponding patient information. 4) Data embedding: To embed the digital envelope into the image. This is done by replacing the least significant bit of a random pixel of the mammogram by one bit of the digital envelope bit stream and repeating for all bits in the bit stream. Experiments with digital IMs demonstrate the following. 1) In the expert center, only the user who knows the private key can open the digital envelope and read the patient information data and the digital signature of the mammogram transmitted from the examination site. 2) Data integrity can be verified by matching the image hash value decrypted from the digital signature with that computed from the transmitted image. 3) No visual quality degradation is detected in the embedded image compared with the original. Our preliminary results demonstrate that AIDM is an effective method for image authenticity and integrity in telemammography application.
Xiaoqiang Zhou, H. K. Huang, Shieh-Liang Lou
IEEE Trans. Medical Imaging1
2000 Real-time teleconsultation with high-resolution and large-volume medical images for collaborative healthcare
abstract
Real-time consultation between referring physicians or radiologists with an expert is critical for timely and adequate management of problem cases. During consultation, both sides need to: 1) synchronously manipulate high-resolution digital radiographic images or large volume MR/CT images; 2) perform interpretation interactively; and 3) converse with audio. We present a specifically designed teleconsultation system in a digital imaging and communication in medicine picture archiving and communication systems clinical environment. The system uses bidirectional remote control technology to meet critical teleconsultation application requirements with high-resolution and large-volume medical images operated in a limited-bandwidth network setting. We give the system design and implementation methods, and also describe the teleconsultation procedure and protocol used in this system. Finally, laboratory and clinical evaluation results are discussed.
Johannes N. Stahl, H. K. Huang, Xiaoqiang Zhou, Shieh-Liang Lou, K. S. Song
IEEE Trans. Inf. Technol. Biomed.4