Fei Gao 0006

dblp:16/722-6 · DBLP profile ↗
← Back
58ranked-venue papers
15as first author
36since 2021 · last 2026
0000-0002-8714-0975ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 32 since 2021Artificial intelligence and machine learning · 29 · 9 first-author · 16 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
YearPublicationVenuePosition
2026 CO²IF: Language-Bridging Hyperspectral-Multispectral Image Fusion with Coordinated and Cross-modal Optimal Transport
abstract
Due to the difficulties of directly obtaining high-resolution hyperspectral images (HR-HSI), the fusion of low-resolution hyperspectral images (LR-HSI) and high-resolution multispectral images (HR-MSI) has emerged as an effective approach. While existing methods leverage image-level priors from HR-MSI, they often lack explicit semantic guidance for precise detail reconstruction. Recognizing that textual scene descriptions encapsulate valuable object attributes and contextual information, we introduce the first Language-Bridging framework for Hyperspectral and Multispectral image fusion (CO²IF). CO²IF leverages language semantics as prior knowledge to explicitly guide the reconstruction process. To bridge the modality gap between textual descriptions and high-dimensional hyperspectral data, we design a Cross-modal Optimal Transport (COT) module. COT establishes precise semantic correspondences between language features and the visual cues of individual spectral bands. Building upon this semantic alignment, we develop a Multimodal Coordinated State Space Model (CoMamba). CoMamba effectively integrates the language-derived priors with spatial information from HR-MSI and spectral information from LR-HSI. This language-guided reconstruction significantly enhances the extraction of crucial spatial-spectral details, leading to superior fidelity in the generated HR-HSI. In addition, this paper adds text descriptions for three widely used datasets. Both qualitative and quantitative experimental results on the public datasets confirm the superiority of the proposed method compared to the SOTA methods.
Mingjin Zhang, Zhongkai Yang, Fei Gao 0006
AAAI3
2026 AdaNoise: Cycle-Consistent Image Translation With Domain-Adaptive Noise Perturbation
abstract
Image-to-image (I2I) translation aims to transform an input image into a target domain while preserving its structural details. Recent advances in diffusion-based generative models have significantly improved the perceptual quality of generated images; however, these approaches still face challenges in controllability and consistency, largely due to the inherent randomness introduced by stochastic noise during the generation process. Specifically, directly manipulating noise distributions without semantic alignment can lead to mode collapse, texture distortion, or loss of domain-specific features. To overcome these challenges, we propose the Adaptive Noise Framework (AdaNoise), a novel and cycle-consistent I2I translation approach guided by domain-adaptive noise modulation. AdaNoise introduces a Domain-Adaptive Noise Perturbation (DANP) module, which adaptively learns structured noise patterns aligned with the target domain distribution, enhancing both the expressiveness and reliability of the translation process. Through integration with a Cycle-Consistent Dual Diffusion (CDD) architecture, AdaNoise ensures faithful content reconstruction while allowing semantically meaningful domain shifts. The framework is designed to maintain a balance between generation quality and controllability, enabling more faithful and flexible image translation. Extensive experiments on two tasks, including SAR-to-optical image translation and low-light image enhancement, validate that AdaNoise not only surpasses existing state-of-the-art methods in terms of image fidelity and semantic preservation, but also achieves more controllable and diverse outputs across varying conditions, thus offering a robust and scalable solution for cross-domain visual generation.
Xi Yang 0011, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 QuARF: Quality-Adaptive Receptive Fields for Degraded Image Perception
abstract
Advanced Deep Neural Networks (DNNs) perform well for high-quality images, but their performance dramatically decreases for degraded images. Data augmentation is commonly used to alleviate this problem, but using too much perturbed data might seriously decrease the performance on pristine images. To tackle this challenge, we take our cue from the assumption of spatial coincidence in human visual perception, i.e. multiscale and varying receptive fields are required for understanding pristine and degraded images. Correspondingly, we propose a novel plug-and-play network architecture, dubbed Quality-Adaptive Receptive Fields (QuARF), to automatically select the optimal receptive fields based on the quality of the input image. To this end, we first design a multi-kernel convolutional block, which comprises multiscale continuous receptive fields. Afterward, we design a quality-adaptive routing network to predict the significance of each kernel, based on the quality features extracted from the input image. In this way, QuARF automatically selects the optimal inference route for each image. To further boost efficiency and effectiveness, the input feature map is split into multiple groups, with each group independently learning its quality-adaptive routing parameters. We apply QuARF to a variety of DNNs and conduct experiments in both discriminative and generation tasks, including semantic segmentation, image translation, and restoration. Thorough experimental results show that QuARF significantly and robustly improves the performance for degraded images, and outperforms data augmentation in most cases.
Fei Gao 0006, Ziyun Li 0002, Wenwang Han, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
AAAI1
2025 IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) focuses on identifying small targets in infrared images. Despite advancements with deep learning, challenges persist due to the IR long-range imaging mechanism, where targets are small, dim, and easily lost in noise and background clutter. Current deep learning methods struggle to suppress noise and background interference while preserving fine details, leading to missed detections and false alarms. To address these issues, we propose IRMamba, an encoder-decoder architecture featuring Pixel Difference Mamba (PDMamba) and a Layer Restoration Module (LRM). Specifically, PDMamba integrates the intensity and directional information of pixel differences between scanning positions and their central neighborhoods into the state equation of the state space model (SSM). This enhances target detail representation and suppresses background interference by capturing local 2D dependencies from a global perspective. In addition, LRM incorporates the double-depth image prior into the iterative convergence algorithm, and utilizes the inter-layer interrelationships to gradually reverse the separation of the target layer, achieving noise suppression and refined reconstruction of the image mask. Experiments conducted on multiple public datasets, including NUAA-SIRST, NUDT-SIRST, and IRSTD-1K, demonstrate the significant advantages of IRMamba over SOTA methods.
Mingjin Zhang, Fei Gao 0006, Jie Guo 0009
AAAI3
2025 MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection
abstract
In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available.
Mingjin Zhang, Yuanjun Ouyang, Fei Gao 0006, Jie Guo 0009, Qiming Zhang 0001, Jing Zhang 0037
AAAI3
2025 Semi-supervised Infrared Small Target Detection with Thermodynamic-Inspired Uneven Perturbation and Confidence Adaptation
abstract
Single-frame Infrared Small Target (SIRST) detection has made significant advancements, but it still faces challenges due to limited labeled data and the foreground-background class imbalance. To address these issues, we introduce a novel Semi-Supervised SIRST Detection (S^3D) pipeline in this paper. First, drawing inspiration from thermodynamics, we propose augmenting infrared images using both chromatically and spatially uneven perturbations. This dual-stream perturbation enhances the diversity and balance of infrared samples, contributing to the robustness of detection models. Additionally, we develop a confidence-adaptive matching method to maintain weighted consistency among perturbed unlabeled samples. Second, to tackle class imbalance in labeled data, we compel the model to generate discriminative predictions for challenging, misclassified examples while down-weighting well-classified examples. We achieve this by modifying the standard cross-entropy loss to squeeze the detector and truncating the loss on well-classified examples. Our innovative Truncated Squeeze (TS) loss focuses on learning discriminative representations for difficult cases and prevents over-optimization for simpler ones. To assess the effectiveness of the perturbation techniques and loss functions, we apply them to various SIRST detectors and conduct comprehensive experiments on two benchmark datasets. Notably, our proposed methods consistently and significantly improve accuracy. Remarkably, our approach achieves over 98% performance of the state-of-the-art fully-supervised method using only 1/8 of the labeled samples.
Mingjin Zhang, Wenteng Shang, Fei Gao 0006, Qiming Zhang 0001, Fengqin Lu, Jing Zhang 0037
AAAI3
2025 SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretraining
abstract
Infrared Small Target Detection (IRSTD) aims to identify low signal-to-noise ratio small targets in infrared images with complex backgrounds, which is crucial for various applications. However, existing IRSTD methods typically rely solely on image modalities for processing, which fail to fully capture contextual information, leading to limited detection accuracy and adaptability in complex environments. Inspired by vision-language models, this paper proposes a novel framework, SAIST, which integrates textual information with image modalities to enhance IRSTD performance. The framework consists of two main components: Scene Recognition Contrastive Language-Image Pretraining (SR-CLIP) and CLIP-guided Segment Anything Model (CG-SAM). SR-CLIP generates a set of visual descriptions through object-object similarity and object-scene relevance, embedding them into learnable prompts to refine the textual description set. This reduces the domain gap between vision and language, generating precise textual and visual prompts. CG-SAM utilizes the prompts generated by SR-CLIP to accurately guide the Mask Decoder in learning prior knowledge of background features, while incorporating infrared imaging equations to improve small target recognition in complex backgrounds and significantly reduce the false alarm rate. Additionally, this paper introduces the first multimodal IRSTD dataset, MIRSTD, which contains abundant image-text pairs. Experimental results demonstrate that the proposed SAIST method outperforms existing state-of-the-art approaches.
Mingjin Zhang, Fei Gao 0006, Jie Guo 0009, Xinbo Gao 0001, Jing Zhang 0037
CVPR3
2025 MAJoR: Visual Emotion Analysis via Multi-Attribute Joint Reasoning
abstract
Visual Emotion Analysis (VEA) seeks to anticipate individuals’ emotional reactions to visual stimuli. The subjective perception of visual emotion is an integrated impact of the appearance, scene, and objects presented in an image. It is thus significance to analysis visual emotion by incorporating diverse visual attributes. Motivated by this, in this paper, we propose a novel VEA method based on Multi-Attribute Joint Reasoning (MAJoR). Specifically, we first use a multi-stream networks to learning multi-attribute representations, including the color, brightness, scene type, and object class of an input image. Afterward, we use a Graph Convolution Network (GCN) to model the inherent relationships among such visual attributes, and to predict the emotion category. Finally, we propose a two-stage knowledge distillation strategy, to boost the performance of light-weight VEA models via MAJoR. Extensive experiments conducted on several VEA databases showcase the superiority of the proposed MAJoR model and the distilled lightweight versions, compared to state-of-the-art approaches. Our code and models are available at: https://github.com/AiArt-Gao/MAJoR.
Yuxin Fei, Jinlan Xu, Maoying Qiao, Fei Gao 0006
ICASSP4
2025 Q-Norm: Robust Representation Learning via Quality-Adaptive Normalization
Lanning Zhang, Fei Gao 0006, Ziyun Li 0002, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
ICCV3
2025 Mixture-of-Modality-Experts for Unified Image Aesthetic Assessment with Multi-Level Adaptation
abstract
Multi-modal image aesthetic assessment (MIAA) has gained significant progress, by predicting aesthetic based on both an image and its text comments. However, most MIAA methods are not applicable, when there are no text comments available. To combat this challenge, we propose a unified image aesthetic assessment (IAA) framework, termed AesFormer, by using mixtures of vision-language Transformers. Specially, AesFormer first learns aligned image-text representations through contrastive learning, and uses a vision-language head for MIAA prediction. Afterward, we propose a multi-level adaptation (MLA) method to adapt the learned MIAA model to the case without text comments, and use another vision head for vison-only IAA (VIAA) prediction. Extensive experimental results show that AesFormer significantly outperforms previous methods in both MIAA and VIAA tasks, on diverse benchmarking datasets. Our code has been released at: https://github.com/AiArt-Gao/AesFormer
Fei Gao 0006, Xiaodan Zhang 0005, Lihuo He, Nannan Wang 0001
ICME1
2025 Towards Aligned Data Forgetting via Twin Machine Unlearning
abstract
Modern privacy regulations have spurred the evolution of machine unlearning, a technique enabling a trained model to efficiently forget specific training data. In prior unlearning methods, the concept of “data forgetting” is often interpreted and implemented as achieving zero classification accuracy on such data. Nevertheless, the authentic aim of machine unlearning is to achieve alignment between the unlearned model and the gold model, i.e., encouraging them to have identical classification accuracy. On the other hand, the gold model often exhibits non-zero classification accuracy due to its generalization ability. To achieve aligned data forgetting, we propose a Twin Machine Unlearning (TMU) approach, where a twin unlearning problem is defined corresponding to the original unlearning problem. Consequently, the generalization-label predictor trained on the twin problem can be transferred to the original problem, facilitating aligned data forgetting. Comprehensive empirical experiments illustrate that our approach significantly enhances the alignment between the unlearned model and the gold model.
Haoxuan Ji, Yuyao Sun, Fei Gao 0006, Haichang Gao, Zhenxing Niu
ICME4
2025 BI-RADS Boosted Breast Cancer Diagnosis With Masked Pretraining On Imbalanced Ultrasound Data
abstract
In clinical diagnosis, the Breast Imaging Reporting and Data System (BI-RADS) levels are highly correlated with pathological categories (benign or malignant). Thus, in this paper, we propose a BI-RADS Boosted Breast Cancer Diagnosis (B3CD) method, for joint predicting both the BI-RADS levels and pathological categories. Specifically, we first train two networks for each task for learning task specific features, and then fuse them through dual spatial attention. Besides, the network backbones are initialized through masked pretraining, due to the limited amount of labeled data. A balanced cross-entropy loss is used for the BI-RADS prediction branch to combat the extremely imbalanced distribution of BI-RADS levels. Experimental results demonstrate that B3CD achieves remarkably superior performance in both breast cancer diagnosis and BI-RADS prediction tasks, across the GDPH&SYSUCC, BUSBRA, and Breast-Lesions-USG datasets. Our code has been released at: https://github.com/AiArt-Gao/B3CD.
Xueqian Pang, Ziyun Li 0002, Junhui Lv, Ruiquan Ge, Zhuoxuan Wu, Fei Gao 0006
ICME6
2025 ReCLIP: Reconstruction-Refined Zero-/Few-Shot Anomaly Classification and Segmentation
abstract
Recent advancements in zero-/few-shot anomaly detection have demonstrated the efficiency of contrastive learning approaches. However, existing methods struggle with imprecise perception of anomaly details and often lack focus on identifying anomaly types. To address these challenges, we propose a flexible reconstruction-refined framework based on contrastive learning, which comprises three core components: a cross-modal alignment network, a reconstruction module, and a dual-attention refinement module. The reconstruction module captures fine-grained anomaly embeddings and fidelity scores, enabling flexible module switching based on task requirements. The attention module directs the cross-modal alignment network to focus on fine-grained anomaly information for accurate segmentation. Our framework leverages the strengths of current anomaly detection algorithms, and excels in zero-/few-shot tasks across industrial datasets, recognizing 43 anomaly types without separate training for each category, and significantly outperforms existing methods, particularly on the challenging MPDD dataset. Our code has been released at https://github.com/AiArt-Gao/ReCLIP.
Lanning Zhang, Yali Shi, Shujie Lan, Fei Gao 0006, Hao Qin 0001, Nannan Wang 0001
ICME4
2025 Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive Imaging
abstract
The objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information for accurate reconstruction. Recognizing that textual description contain rich semantic information that can significantly enhance details, this paper introduces a novel framework, CAMM, which integrates text information into the model to improve the performance. The framework comprises two key components: Fine-grained Alignment Module (FAM) and Multimodal Fusion Mamba (MFM). Specifically, FAM is used to reduce the knowledge gap between the RGB domain obtained by the pre-trained vision-language model and the HSI domain. Through the double constraints of distribution similarity and entropy, the adaptive alignment of different complexity features is realized, which makes the encoded features more accurate. MFM aims to identify the guiding effect of RGB features and text features on HSI in space and channel dimensions. Instead of fusing features directly, it integrates prior at image-level and text-level prior into Mamba's state-space equation, so that each scanning step can be accurately guided. This kind of positive feedback adjustment ensures the authenticity of the guiding information. To our knowledge, this is the first text-guided model for compressive spectral imaging. Extensive experimental results the public datasets demonstrate the superior performance of CAMM, validating the effectiveness of our proposed method.
Mingjin Zhang, Longyi Li, Fei Gao 0006, Qiming Zhang 0001, Jie Guo 0009
IJCAI3
2025 CADQ: Attribute-Consistent Face Cartoonization with Cross-modal Aligned and Deformable Quantization
abstract
Face cartoonization remains a challenging task due to significant geometric deformations between facial photos and cartoons, as well as the absence of paired training data for supervised learning. Existing methods struggle to generate high-quality cartoonized avatars with attribute consistency. To address this challenge, this paper proposes an unsupervised facial cartoonization method based on cross-domain aligned and deformable vector quantization (CADQ). Firstly, we construct textual descriptions with facial attributes for both photo datasets and cartoon collections. Attribute consistency during transformation is enforced through individually contrastive learning between image-text cross-modal features and globally distribution alignment across photo-cartoon domains. Secondly, a deformable Transformer with dual attention is introduced during the transformation process, which queries corresponding cartoon codebook entries based on image features to simulate cross-domain geometric deformations. Experimental results demonstrate that the proposed method can convert facial photos into high-quality cartoons with attribute consistency, outperforming existing state-of-the-art approaches. Furthermore, the method can be effectively extended to unsupervised cross-domain generation of other artistic portrait styles, achieving superior or highly competitive performance. Our code has been released at: https://github.com/IIP-Lab-XDU/CADQ.
Yongjie Hu, Ziyun Li 0002, Fei Gao 0006, Henrik Boström, Nannan Wang 0001
ACM Multimedia4
2025 S3OIL: Semi-Supervised SAR-to-Optical Image Translation via Multi-Scale and Cross-Set Matching
abstract
Image-to-image translation has achieved great success, but still faces the significant challenge of limited paired data, particularly in translatingSynthetic Aperture Radar(SAR) images to optical images. Furthermore, most existing semi-supervised methods place limited emphasis on leveraging the data distribution. To address those challenges, we propose aSemi-Supervised SAR-to-Optical Image Translation(S3OIL) method that achieves high-quality image generation using minimal paired data and extensive unpaired data while strategically exploiting the data distribution. To this end, we first introduce aCross-Set Alignment Matching(CAM) mechanism to create local correspondences between the generated results of paired and unpaired data, ensuring cross-set consistency. In addition, for unpaired data, we apply weak and strong perturbations and establish intra-setMulti-Scale Matching(MSM) constraints. For paired data, intra-modal semantic consistency (ISC) is presented to ensure alignment with the ground truth. Finally, we propose local and global cross-modal semantic consistency (CSC) to boost structural identity during translation. We conduct extensive experiments on SAR-to-optical datasets and another sketch-to-anime task, demonstrating that S3OIL delivers competitive performance compared to state-of-the-art unsupervised, supervised, and semi-supervised methods, both quantitatively and qualitatively. Ablation studies further reveal that S3OIL can ensure the preservation of both semantic content and structural integrity of the generated images. Our code is available at: https://github.com/XduShi/SOIL.
Xi Yang 0011, Ziyun Li 0002, Maoying Qiao, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Image Process.5
2025 Boosting Modal-Specific Representations for Sentiment Analysis With Incomplete Modalities
abstract
Multimodal sentiment analysis aims at exploiting complementary information from multiple modalities or data sources to enhance the understanding and interpretation of sentiment. While existing multi-modal fusion techniques offer significant improvements in sentiment analysis, real-world scenarios often involve missing modalities, introducing complexity due to uncertainty of which modalities may be absent. To tackle the challenge of incomplete modality-specific feature extraction caused by missing modalities, this paper proposes a Cosine Margin-Aware Network (CMANet) which centers on the Cosine Margin-Aware Distillation (CMAD) module. The core module measures distance between samples and the classification boundary, enabling CMANet to focus on samples near the boundary. So, it effectively captures the unique features of different modal combinations. To address the issue of modality imbalance during modality-specific feature extraction, this paper proposes a Weak Modality Regularization (WMR) strategy, which aligns the feature distributions between strong and weak modalities at the dataset-level, while also enhancing the prediction loss of samples at the sample-level. This dual mechanism improves the recognition robustness of weak modality combination. Extensive experiments demonstrate that the proposed method outperforms the previous best model, MMIN, with a 3.82% improvement in unweighted accuracy. These results underscore the robustness of the approach under conditions of uncertain and missing modalities.
Lihuo He, Fei Gao 0006, Kaifan Zhang, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.3
2024 Generating Handwritten Mathematical Expressions From Symbol Graphs: An End-to-End Pipeline
abstract
In this paper, we explore a novel challenging generation task, i.e. Handwritten Mathematical Expression Generation (HMEG) from symbolic sequences. Since symbolic sequences are naturally graph-structured data, we formulate HMEG as a graph-to-image (G2I) generation problem. Unlike the generation of natural images, HMEG requires critic layout clarity for synthesizing correct and recognizable formulas, but has no real masks available to supervise the learning process. To alleviate this challenge, we propose a novel end-to-end G2I generation pipeline (i.e. graph → layout →mask →image), which requires no real masks or nondifferentiable alignment between layouts and masks. Technically, to boost the capacity of predicting detailed relations among adjacent symbols, we propose a Less-is-More (LiM) learning strategy. In addition, we design a differentiable layout refinement module, which maps bounding boxes to pixel-level soft masks, so as to further alleviate ambiguous layout areas. Our whole model, including layout prediction, mask refinement, and image generation, can be jointly optimized in an end-to-end manner. Experimental results show that, our model can generate highquality HME images, and outperforms previous generative methods. Besides, a series of ablations study demonstrate effectiveness of the proposed techniques. Finally, we validate that our generated images promisingly boosts the performance of HME recognition models, through data augmentation. Our code and results are available at: https://github.com/AiArt-HDU/HMEG.
Yu Chen 0003, Fei Gao 0006, Yanguang Zhang, Maoying Qiao, Nannan Wang 0001
CVPR2
2024 Chinese Ink Cartoons Generation With Multiscale Semantic Consistency
abstract
Chinese Ink Cartoon (CIC) generation is challenging, due to the lack of paired photo-cartoon data and the severe geometric deformations between photos and cartoons. To tackle this challenge, in this paper, we build a high-quality CIC dataset with rich annotations, and propose a novel CIC generation method based on Multiscale Semantic Consistency (MuSeC). The CIC dataset consists of $\mathbf{1, 2 0 0}$ high-resolution CIC paintings with nearly 2,000 annotations of human faces and bodies, respectively. The generation task is formulated as an unsupervised image-to-image translation problem. First, we constrain the generated cartoon to pixel-wisely convey the semantic structure of the input photo through a learned CIC face parsing network. Additionally, we use patch-wise contrastive learning and global semantic consistency loss. In this way, the generated cartoon is optimized to precisely present the identity, attributes and structure of the input photo. Experimental results show that the proposed method can generate high-quality CICs, and outperforms previous methods both quantitatively and qualitatively. Our dataset and code have been released at github.com/AiArt-Gao/MuSeC
Shuran Su, Mei Du, Jintao Mao, Renshu Gu, Fei Gao 0006
CW5
2024 Learning Discriminative Style Representations for Unsupervised and Few-Shot Artistic Portrait Drawing Generation
abstract
In this paper, we propose an unsupervised artistic portrait drawing generation method for few-shot datasets based on contrastive learning of style features. Firstly, we construct a discriminative style encoder with contrastive learning, improving the ability of the encoder to separate style features. Secondly, based on the dynamic codebook and momentum network, we used historical average features instead of batch instance features to prevent the problem of style bias in few-shot datasets. Finally, a conditional projection discriminator with filter response normalization is utilized to improve the discriminative ability of the discriminator and the stability of the generative adversarial network, which motivates the generator to synthesize more realistic image details. Quantitative and qualitative analysis show that the method proposed in this paper significantly improves the quality of artistic portrait drawing generation, and outperforms existing benchmarks in terms of visual effect and metrics evaluation. Our code and results are avilable at https://github.com/AiArt-HDU/Co-GAN.
Junkai Fang, Maoying Qiao, Fei Gao 0006
ICASSP6
2024 Human-Robot Interactive Creation of Artistic Portrait Drawings
abstract
In this paper, we present a novel system for Human-Robot Interactive Creation of Artworks (HRICA). Different from previous robot painters, HRICA allows a human user and a robot to alternately draw strokes on a canvas, to collaboratively create a portrait drawing through frequent interactions. The key is to enable the robot to understand human intentions, during the interactive creation process. We here formulate this as a mask-free image inpainting problem, and propose a novel method to estimate the complete version of a portrait drawing, after the human user has drawn some initial strokes. In this way, the robot can select some complementary strokes and draw them on the canvas. To train and evaluate our inpainting method, we construct a novel large-scale portrait drawing dataset, CelebLine, which composes of high-quality portrait line-drawings, with dense labels of both 2D semantic parsing masks and 3D depth maps. Finally, we develop a human-robot interactive drawing system with low-cost hardware, user-friendly interface, and interesting creation experience. Experiments show that our robot can stably cooperate with human users to create diverse styles of portrait drawings. In addition, our portrait drawing inpainting method significantly outperforms previous advanced methods. The code and dataset have been released at: https://github.com/fei-aiart/HRICA.
Fei Gao 0006, Lingna Dai, Jingjie Zhu, Mei Du, Maoying Qiao, Chenghao Xia, Nannan Wang 0001, Peng Li 0031
ICRA1
2024 AesMamba: Universal Image Aesthetic Assessment with State Space Models
abstract
Image Aesthetic Assessment (IAA) aims to objectively predict the generic or personalized evaluations, of the aesthetic or fine-grained multi-attributes, based on visual or multimodal inputs. Previously, researchers have designed diverse and specialized methods, for specific IAA tasks, based on different input-output situations. Is it possible to design a universal IAA framework applicable for the whole IAA task taxonomy? In this paper, we explore this issue, and propose a modular IAA framework, dubbed AesMamba. Specially, we use the Visual State Space Model (VMamba), instead of CNNs or ViTs, to learn comprehensive representations of aesthetic-related attributes; because VMamba can efficiently achieve both global and local effective receptive fields. Afterward, a modal-adaptive module is used to automatically produce the integrated representations, conditioned on the type of input. In the prediction module, we propose a Multitask Balanced Adaptation (MBA) module, to boost task-specific features, with emphasis on the tail instances. Finally, we formulate the personalized IAA task as a multimodal learning problem, by converting a user's anonymous subject characters to a text prompt. This prompting strategy effectively employs the semantics of flexibly selected characters, for inferring individual preferences. AesMamba can be applied to diverse IAA tasks, through flexible combination of these modules. Extensive experiments on numerous datasets, demonstrate that AesMamba consistently achieves superior or competitive performance, on all IAA tasks, in comparison with previous SOTA methods. The code has been released at https://github.com/AiArt-Gao/AesMamba Github.
Fei Gao 0006, Maoying Qiao, Nannan Wang 0001
ACM Multimedia1
2024 Micro-Action Recognition via Hierarchical Fusion and Inference
abstract
Micro-actions are spontaneous body movements that indicate a person's true feelings and potential intentions, and micro-action recognition is important in human behavior analysis. Yet, recognizing micro-actions is challenging because they are subtle and appear for a very short time compared to normal actions. In this paper, we propose a micro-action recognition framework based on Hierarchical Fusion and Inference (HiFI) to capture subtle multimodal information. Specifically, we first hierarchically integrate multimodal local and global information, including the 2D key-points of faces, hands and bodies, the depth information, and the RGB image sequences. Afterward, both 3D-CNNs and Transformers are used to effectively capture local and long-range dependence. Finally, we propose a novel from-fine-to-coarse (F2C) inference strategy, based on hybrid ensemble of multi-branches, to boost the accuracy and credibility of coarse action recognition. Our solution ranked 4th in the MAC Challenge Track 1.
Fan Gong, Qijian Bao, Fei Gao 0006, Renshu Gu, Gang Xu 0001
ACM Multimedia5
2024 3D Human Pose Estimation from Multiple Dynamic Views via Single-view Pretraining with Procrustes Alignment
abstract
3D Human pose estimation from multiple cameras with unknown calibration has received less attention than it should. The few existing data-driven solutions do not fully exploit 3D training data that are available on the market, and typically train from scratch for every novel multi-view scene, which impedes both accuracy and efficiency. We show how to exploit 3D training data to the fullest and associate multiple dynamic views efficiently to achieve high precision on novel scenes using a simple yet effective framework, dubbed Multiple Dynamic View Pose estimation (MDVPose). MDVPose utilizes novel scenarios data to finetune a single-view pretrained motion encoder in multi-view setting, aligns arbitrary number of views in a unified coordinate via Procruste alignment, and imposes multi-view consistency. The proposed method achieves 22.1 mm P-MPJPE or 34.2 mm MPJPE on the challenging in-the-wild Ski-Pose PTZ dataset, which outperforms the state-of-the-art method by 24.8% P-MPJPE (-7.3 mm) and 19.0% MPJPE (-8.0 mm). It also outperforms the state-of-the-art methods by a large margin (-18.2mm P-MPJPE and -28.3mm MPJPE) on the EgoBody dataset. In addition, MDVPose achieves robust performance on the Human3.6M datasets featuring multiple static cameras. Code is available at https://github.com/iGame-Lab/MDVPose.
Renshu Gu, Yixuan Si, Fei Gao 0006, Jiamin Xu, Gang Xu 0001
ACM Multimedia4
2023 Masked and Adaptive Transformer for Exemplar Based Image Translation
abstract
We present a novel framework for exemplar based image translation. Recent advanced methods for this task mainly focus on establishing cross-domain semantic correspondence, which sequentially dominates image generation in the manner of local style control. Unfortunately, cross-domain semantic matching is challenging; and matching errors ultimately degrade the quality of generated images. To overcome this challenge, we improve the accuracy of matching on the one hand, and diminish the role of matching in image generation on the other hand. To achieve the former, we propose a masked and adaptive transformer (MAT) for learning accurate cross-domain correspondence, and executing context-aware feature augmentation. To achieve the latter, we use source features of the input and global style codes of the exemplar, as sup-plementary information, for decoding an image. Besides, we devise a novel contrastive style learning method, for acquire quality-discriminative style representations, which in turn benefit high-quality image generation. Experimen-tal results show that our method, dubbed MATEBIT, performs considerably better than state-of-the-art methods, in diverse image translation tasks. The codes are available at https://github.com/AiArt-HDU/MATEBIT.
Fei Gao 0006, Nannan Wang 0001, Gang Xu 0001
CVPR2
2023 Human-Inspired Facial Sketch Synthesis with Dynamic Adaptation
abstract
Facial sketch synthesis (FSS) aims to generate a vivid sketch portrait from a given facial photo. Existing FSS methods merely rely on 2D representations of facial semantic or appearance. However, professional human artists usually use outlines or shadings to covey 3D geometry. Thus facial 3D geometry (e.g. depth map) is extremely important for FSS. Besides, different artists may use diverse drawing techniques and create multiple styles of sketches; but the style is globally consistent in a sketch. Inspired by such observations, in this paper, we propose a novel Human-Inspired Dynamic Adaptation (HIDA) method. Specially, we propose to dynamically modulate neuron activations based on a joint consideration of both facial 3D geometry and 2D appearance, as well as globally consistent style control. Besides, we use deformable convolutions at coarse-scales to align deep features, for generating abstract and distinct outlines. Experiments show that HIDA can generate high-quality sketches in multiple styles, and significantly outperforms previous methods, over a large range of challenging faces. Besides, HIDA allows precise style control of the synthesized sketch, and generalizes well to natural scenes and other artistic styles. Our code and results have been released online at: https://github.com/AiArt-HDU/HIDA.
Fei Gao 0006, Nannan Wang 0001
ICCV1
2023 Semantic-Aware Generation of Multi-View Portrait Drawings
abstract
Neural radiance fields (NeRF) based methods have shown amazing performance in synthesizing 3D-consistent photographic images, but fail to generate multi-view portrait drawings. The key is that the basic assumption of these methods -- a surface point is consistent when rendered from different views -- doesn't hold for drawings. In a portrait drawing, the appearance of a facial point may changes when viewed from different angles. Besides, portrait drawings usually present little 3D information and suffer from insufficient training data. To combat this challenge, in this paper, we propose a Semantic-Aware GEnerator (SAGE) for synthesizing multi-view portrait drawings. Our motivation is that facial semantic labels are view-consistent and correlate with drawing techniques. We therefore propose to collaboratively synthesize multi-view semantic maps and the corresponding portrait drawings. To facilitate training, we design a semantic-aware domain translator, which generates portrait drawings based on features of photographic faces. In addition, use data augmentation via synthesis to mitigate collapsed results. We apply SAGE to synthesize multi-view portrait drawings in diverse artistic styles. Experimental results show that SAGE achieves significantly superior or highly competitive performance, compared to existing 3D-aware image synthesis methods. The codes are available at https://github.com/AiArt-HDU/SAGE.
Fei Gao 0006, Nannan Wang 0001, Gang Xu 0001
IJCAI2
2023 Explicitly Semantic Guidance for Face Sketch Attribute Recognition With Imbalanced Data
abstract
Current facial attribute recognition (FAR) methods focus exclusively on photographs, and fail when applied to face sketches. Besides, face sketch attribute recognition (FSAR) encounters the following difficulties: the scarcity of labelled instances, the heavily imbalanced data distribution, and the inter-attribute correlations. To combat this challenge, in this paper, we propose a novel FSAR method based on the correlations between facial attributes and semantic regions. Our full model includes a shared feature extraction network, followed by several attribute-specific prediction branches. In each branch, we use the corresponding semantic mask, to select features from the associated region, for attribute prediction. Such explicitly semantic guidance (ESG) reduces the learning space, and thus alleviates the problems of limited data and imbalanced distribution. Besides, ESG decouples inter-attribute correlations, and makes the recognition process credible. Finally, we adopt the balanced cross-entropy loss during training, which further alleviates the problem of imbalanced data distribution. Experiments on the benchmark FS2K dataset demonstrate that our method significantly outperforms advanced visual recognition networks. Our codes have been released at:https://github.com/AiArt-HDU/ESGAR.
Shahadat Shahed, Jiangnan Hong, Fei Gao 0006
IEEE Signal Process. Lett.5
2022 Sample-Efficient Kernel Mean Estimator with Marginalized Corrupted Data
abstract
Estimating the kernel mean in a reproducing kernel Hilbert space is central to many kernel-based learning algorithms. Given a finite sample, an empirical average is used as a standard estimation of the target kernel mean. Prior works have shown that better estimators can be constructed by shrinkage methods. In this work, we propose to corrupt data examples with noise from known distributions and present a new kernel mean estimator, called the marginalized kernel mean estimator, which estimates kernel mean under the corrupted distributions. Theoretically, we justify that the marginalized kernel mean estimator introduces implicit regularization in kernel mean estimation. Empirically, on a variety of tasks, we show that the marginalized kernel mean estimator is sample-efficient and obtains much lower estimation errors than the existing estimators.
Xiaobo Xia, Mingming Gong, Nannan Wang 0001, Fei Gao 0006, Haikun Wei, Tongliang Liu
KDD5
2022 IGA-Reuse-NET: A deep-learning-based isogeometric analysis-reuse approach with topology-consistent parameterization
abstract
In this paper, a deep learning framework combined with isogeometric analysis (IGA for short) called IGA-Reuse-Net is proposed for efficient reuse of numerical simulation on a set of topology-consistent models. Compared with previous data-driven numerical simulation methods only for simple computational domains, our method can predict high-accuracy PDE solutions over topology-consistent geometries with complex boundaries. UNet3+ architecture with interlaced sparse self-attention (ISSA) module is used to enhance the performance of the network. In addition, we propose a new loss function that combines a coefficients loss and a numerical solution loss. Several training datasets with topology-consistent models are constructed for the proposed framework. To verify the effectiveness of our approach, two different types of Poisson equations with different source functions are solved on three datasets with different topologies. Our framework can achieve a good trade-off between accuracy and efficiency. It outperforms the physics-informed neural network (PINN for short) model and yields promising results of prediction.
Jinlan Xu, Fei Gao 0006, Charlie C. L. Wang, Renshu Gu, Timon Rabczuk, Gang Xu 0001
Comput. Aided Geom. Des.3
2021 Bridging Unpaired Facial Photos and Sketches by Line-Drawings
abstract
In this paper, we propose a novel method to learn face sketch synthesis models by using unpaired data. Our main idea is bridging the photo domain ${\mathcal{X}}$ and the sketch domain Y by using the line-drawing domain ${\mathcal{Z}}$. Specially, we map both photos and sketches to line-drawings by using a neural style transfer method, i.e. $F:{\mathcal{X}}/{\mathcal{Y}} \mapsto {\mathcal{Z}}$. Consequently, we obtain pseudo paired data $({\mathcal{Z}},{\mathcal{Y}})$, and can learn the mapping $G:{\mathcal{Z}} \mapsto {\mathcal{Y}}$ in a supervised learning manner. In the inference stage, given a facial photo, we can first transfer it to a line-drawing and then to a sketch by G ○ F. Additionally, we propose a novel stroke loss for generating different types of strokes. Our method, termed sRender, accords well with human artists’ rendering process. Experimental results demonstrate that sRender can generate multi-style sketches, and significantly outperforms existing unpaired image-to-image translation methods.
Meimei Shang, Fei Gao 0006, Xiang Li 0205, Jingjie Zhu, Lingna Dai
ICASSP2
2021 High-Quality Face Sketch Synthesis via Geometric Normalization and Regularization
abstract
In this work, we propose a novel Generative Adversarial Network for generating a structure-consistent and texture-realistic sketch, conditioned on a face photo. To this end, we propose to boost the capacity of the generator via geometric normalization and regularization. Specially, we first propose an enhanced spatially-adaptive normalization module to modulate the activation, based on the semantic layout and encoding features of the input face. Besides, we use two regularization loss functions to minimize the structural divergence between a generated sketch and the corresponding face photo. Experimental results show that our proposed techniques significantly improve the quality of synthesized sketches, in terms of both structure and texture. Besides, our full model can generate high-quality sketches and significantly outperform previous state-of-the-arts, over a wide range of challenging data. We have made our code and results publicly available: http://aiart.live/genre/.
Xiang Li 0205, Fei Gao 0006
ICME2
2021 Learning efficient, explainable and discriminative representations for pulmonary nodules classification
Hanliang Jiang, Fuhao Shen, Fei Gao 0006, Weidong Han 0001
Pattern Recognit.3
2021 Learning stacking regression for no-reference super-resolution image quality assessment
Kaibing Zhang, Danni Zhu, Jie Li 0001, Xinbo Gao 0001, Fei Gao 0006
Signal Process.5
2021 Toward Realistic Face Photo-Sketch Synthesis via Composition-Aided GANs
abstract
Face photo-sketch synthesis aims at generating a facial sketch/photo conditioned on a given photo/sketch. It covers wide applications including digital entertainment and law enforcement. Precisely depicting face photos/sketches remains challenging due to the restrictions on structural realism and textural consistency. While existing methods achieve compelling results, they mostly yield blurred effects and great deformation over various facial components, leading to the unrealistic feeling of synthesized images. To tackle this challenge, in this article, we propose using facial composition information to help the synthesis of face sketch/photo. Especially, we propose a novel composition-aided generative adversarial network (CA-GAN) for face photo-sketch synthesis. In CA-GAN, we utilize paired inputs, including a face photo/sketch and the corresponding pixelwise face labels for generating a sketch/photo. Next, to focus training on hard-generated components and delicate facial structures, we propose a compositional reconstruction loss. In addition, we employ a perceptual loss function to encourage the synthesized image and real image to be perceptually similar. Finally, we use stacked CA-GANs (SCA-GANs) to further rectify defects and add compelling details. The experimental results show that our method is capable of generating both visually comfortable and identity-preserving face sketches/photos over a wide range of challenging data. In addition, our method significantly decreases the best previous Fréchet inception distance (FID) from 36.2 to 26.2 for sketch synthesis, and from 60.9 to 30.5 for photo synthesis. Besides, we demonstrate that the proposed method is of considerable generalization ability.
Jun Yu 0002, Xingxin Xu, Fei Gao 0006, Shengjie Shi, Meng Wang 0001, Dacheng Tao, Qingming Huang
IEEE Trans. Cybern.3
2021 Complementary, Heterogeneous and Adversarial Networks for Image-to-Image Translation
abstract
Image-to-image translation is to transfer images from a source domain to a target domain. Conditional Generative Adversarial Networks (GANs) have enabled a variety of applications. Initial GANs typically conclude one single generator for generating a target image. Recently, using multiple generators has shown promising results in various tasks. However, generators in these works are typically of homogeneous architectures. In this paper, we argue that heterogeneous generators are complementary to each other and will benefit the generation of images. By heterogeneous, we mean that generators are of different architectures, focus on diverse positions, and perform over multiple scales. To this end, we build two generators by using a deep U-Net and a shallow residual network, respectively. The former concludes a series of down-sampling and up-sampling layers, which typically have large perception field and great spatial locality. In contrast, the residual network has small perceptual fields and works well in characterizing details, especially textures and local patterns. Afterwards, we use a gated fusion network to combine these two generators for producing a final output. The gated fusion unit automatically induces heterogeneous generators to focus on different positions and complement each other. Finally, we propose a novel approach to integrate multi-level and multi-scale features in the discriminator. This multi-layer integration discriminator encourages generators to produce realistic details from coarse to fine scales. We quantitatively and qualitatively evaluate our model on various benchmark datasets. Experimental results demonstrate that our method significantly improves the quality of transferred images, across a variety of image-to-image translation tasks. We have made our code and results publicly available: http://aiart.live/chan/.
Fei Gao 0006, Xingxin Xu, Jun Yu 0002, Meimei Shang, Xiang Li 0205, Dacheng Tao
IEEE Trans. Image Process.1
2020 Making Robots Draw A Vivid Portrait In Two Minutes
abstract
Significant progress has been made with artistic robots. However, existing robots fail to produce high-quality portraits in a short time. In this work, we present a drawing robot, which can automatically transfer a facial picture to a vivid portrait, and then draw it on paper within two minutes averagely. At the heart of our system is a novel portrait synthesis algorithm based on deep learning. Innovatively, we employ a self-consistency loss, which makes the algorithm capable of generating continuous and smooth brush-strokes. Besides, we propose a componential sparsity constraint to reduce the number of brush-strokes over insignificant areas. We also implement a local sketch synthesis algorithm, and several pre- and post-processing techniques to deal with the background and details. The portrait produced by our algorithm successfully captures individual characteristics by using a sparse set of continuous brush-strokes. Finally, the portrait is converted to a sequence of trajectories and reproduced by a 3-degree-of-freedom robotic arm. The whole portrait drawing robotic system is named AiSketcher. Extensive experiments show that AiSketcher can produce considerably high-quality sketches for a wide range of pictures, including faces in-the-wild and universal images of arbitrary content. To our best knowledge, AiSketcher is the first portrait drawing robot that uses neural style transfer techniques. AiSketcher has attended a quite number of exhibitions and shown remarkable performance under diverse circumstances.
Fei Gao 0006, Jingjie Zhu, Zeyuan Yu, Peng Li 0031, Tao Wang 0004
IROS1
2020 Representation learning of image composition for aesthetic prediction
Meimei Shang, Fei Gao 0006, Rongsheng Li, Jun Yu 0002
Comput. Vis. Image Underst.3
2020 Style-adaptive photo aesthetic rating via convolutional neural networks and multi-task learning
Fei Gao 0006, Ziyun Li 0002, Jun Yu 0002, Junze Yu, Qingming Huang, Qi Tian 0001
Neurocomputing1
2020 Attentive and ensemble 3D dual path networks for pulmonary nodules classification
Hanliang Jiang, Fei Gao 0006, Xingxin Xu, Suguo Zhu
Neurocomputing2
2020 Incremental focal loss GANs
Fei Gao 0006, Jingjie Zhu, Hanliang Jiang, Zhenxing Niu, Weidong Han 0001, Jun Yu 0002
Inf. Process. Manag.1
2020 Fashion analysis and understanding with artificial intelligence
Xiaoling Gu, Fei Gao 0006, Min Tan 0005
Inf. Process. Manag.2
2019 Improving Facial Attractiveness Prediction via Co-attention Learning
abstract
Facial attractiveness prediction has drawn considerable attention from image processing community. Despite the substantial progress achieved by existing works, various challenges remain. One is the lack of accurate representation for facial composition, which is essential for attractiveness evaluation. In this paper, we propose to use pixel-wise labelling masks as the meta information of facial composition, and input them into a network for learning high-level semantic representations. The other challenge is to define to what degree different local properties contribute to facial attractiveness. To tackle this challenge, we employ a co-attention learning mechanism to concurrently characterize the significance of different regions and that of distinct facial components. We conduct experiments on the SCUT-FBP5500 and CelebA datasets. Results show that our co-attention learning mechanism significantly improves the facial attractiveness prediction accuracy. Besides, our method consistently produces appealing results and outperforms previous advanced approaches.
Shengjie Shi, Fei Gao 0006, Xuantong Meng, Xingxin Xu, Jingjie Zhu
ICASSP2
2018 Hierarchical convolutional features for end-to-end representation-based visual tracking
Suguo Zhu, Zhenying Fang, Fei Gao 0006
Mach. Vis. Appl.3
2018 Blind image quality prediction by exploiting multi-level deep representations
Fei Gao 0006, Jun Yu 0002, Suguo Zhu, Qingming Huang, Qi Tian 0001
Pattern Recognit.1
2018 Face biometric quality assessment via light CNN
Jun Yu 0002, Kejia Sun, Fei Gao 0006, Suguo Zhu
Pattern Recognit. Lett.3
2018 User-Click-Data-Based Fine-Grained Image Recognition via Weakly Supervised Metric Learning
abstract
We present a novel fine-grained image recognition framework using user click data, which can bridge the semantic gap in distinguishing categories that are similar in visual. As query set in click data is usually large-scale and redundant, we first propose a click-feature-based query-merging approach to merge queries with similar semantics and construct a compact click feature. Afterward, we utilize this compact click feature and convolutional neural network (CNN)-based deep visual feature to jointly represent an image. Finally, with the combined feature, we employ the metriclearning-based template-matching scheme for efficient recognition. Considering the heavy noise in the training data, we introduce a reliability variable to characterize the image reliability, and propose a weakly-supervised metric and template leaning with smooth assumption and click prior (WMTLSC) method to jointly learn the distance metric, object templates, and image reliability. Extensive experiments are conducted on a public Clickture-Dog dataset and our newly established Clickture-Bird dataset. It is shown that the click-data-based query merging helps generating a highly compact (the dimension is reduced to 0.9%) and dense click feature for images, which greatly improves the computational efficiency. Also, introducing this click feature into CNN feature further boosts the recognition accuracy. The proposed framework performs much better than previous state-of-the-arts in fine-grained recognition tasks.
Min Tan 0005, Jun Yu 0002, Zhou Yu 0001, Fei Gao 0006, Yong Rui, Dacheng Tao
ACM Trans. Multim. Comput. Commun. Appl.4
2017 Convolutional neural networks for intestinal hemorrhage detection in wireless capsule endoscopy images
abstract
Wireless capsule endoscopy (WCE) can painlessly capture a large number of images inside the intestine. However, only a small portion of these WCE images contain hemorrhage. It is thus critical to develop automated hemorrhage detection method to facilitate the diagnosis of intestinal diseases. However, automated hemorrhage detection is complicated by 1) the extreme imbalance between the amount of hemorrhage images and that of normal images; and 2) the variety of the appearance, texture, and luminance inside the intestine. In this paper, we proposed to learn a robust intestinal hemorrhage detection model via Convolutional Neural Networks (CNNs), because of CNNs' extraordinary performance in solving various image understanding tasks. Specially, we explored different CNN architectures and data augmentation methods. Besides, we investigated the correlation between hemorrhage detection accuracy and image quality. Across about 1.3k hemorrhage images and 40k normal images, the learned CNN model achieves an F-measure of 98.87%.
Panpeng Li, Ziyun Li 0002, Fei Gao 0006, Jun Yu 0002
ICME3
2017 DeepSim: Deep similarity for image quality assessment
Fei Gao 0006, Panpeng Li, Min Tan 0005, Jun Yu 0002, Yani Zhu
Neurocomputing1
2017 Deep Multimodal Distance Metric Learning Using Click Constraints for Image Ranking
abstract
How do we retrieve images accurately? Also, how do we rank a group of images precisely and efficiently for specific queries? These problems are critical for researchers and engineers to generate a novel image searching engine. First, it is important to obtain an appropriate description that effectively represent the images. In this paper, multimodal features are considered for describing images. The images unique properties are reflected by visual features, which are correlated to each other. However, semantic gaps always exist between images visual features and semantics. Therefore, we utilize click feature to reduce the semantic gap. The second key issue is learning an appropriate distance metric to combine these multimodal features. This paper develops a novel deep multimodal distance metric learning (Deep-MDML) method. A structured ranking model is adopted to utilize both visual and click features in distance metric learning (DML). Specifically, images and their related ranking results are first collected to form the training set. Multimodal features, including click and visual features, are collected with these images. Next, a group of autoencoders is applied to obtain initially a distance metric in different visual spaces, and an MDML method is used to assign optimal weights for different modalities. Next, we conduct alternating optimization to train the ranking model, which is used for the ranking of new queries with click features. Compared with existing image ranking methods, the proposed method adopts a new ranking model to use multimodal features, including click features and visual features in DML. We operated experiments to analyze the proposed Deep-MDML in two benchmark data sets, and the results validate the effects of the method.
Jun Yu 0002, Xiaokang Yang 0001, Fei Gao 0006, Dacheng Tao
IEEE Trans. Cybern.3
2016 Data-driven facial animation via hypergraph learning
abstract
Data-driven facial animation has attracted much attention in recent years. Existing facial animation methods may not preserve the topology structure, and cannot achieve a natural face. This paper proposes a new data-driven facial animation method based on hypergraph learning. It drives a neutral face to a certain expression face. This paper assumes that neutral face has similar topology with the expression face, we compute the alignment laplacian matrix using hypergraph learning. To get a natural face, we add a constraint item which is consisted of a set of motion data. Experiment results demonstrate that our method can achieve a natural expression face. And the results show the superiority over the state-of-art.
Jun Yu 0002, Fei Gao 0006, Jian Zhang 0026
SMC3
2016 Photo aesthetic quality assessment via label distribution learning
abstract
Automatic prediction of photo aesthetic quality is useful for many practical purposes. Current computational approaches typically solved this problem by assigning a categorical label (good or bad) to a photo. However, due to the subjectivity and complexity of humans aesthetic judgments, only a categorical label is insufficient to represent humans perceived aesthetic quality of a photo. This paper focuses on an interesting problem: is it possible to predict the crowed opinions about the aesthetic quality of a photo? The crowed opinion here is expressed by the distribution of scores given by a number of subjects. For each given photo, a deep convolutional neural network (DCNN) is utilized to calculate its feature representation. Afterwards, the crowed opinion prediction problem is formulated as one of label distribution learning (LDL). Experiments show that the proposed method is highly effective and outperforms state-of-the-art algorithms.
Fei Gao 0006, Di Huang 0007, Min Tan 0005, Jun Yu 0002
SMC2
2016 Biologically inspired image quality assessment
Fei Gao 0006, Jun Yu 0002
Signal Process.1
2015 Learning to Rank for Blind Image Quality Assessment
abstract
Blind image quality assessment (BIQA) aims to predict perceptual image quality scores without access to reference images. State-of-the-art BIQA methods typically require subjects to score a large number of images to train a robust model. However, subjective quality scores are imprecise, biased, and inconsistent, and it is challenging to obtain a large-scale database, or to extend existing databases, because of the inconvenience of collecting images, training the subjects, conducting subjective experiments, and realigning human quality evaluations. To combat these limitations, this paper explores and exploits preference image pairs (PIPs) such as the quality of image Ia is better than that of image Ib for training a robust BIQA model. The preference label, representing the relative quality of two images, is generally precise and consistent, and is not sensitive to image content, distortion type, or subject identity; such PIPs can be generated at a very low cost. The proposed BIQA method is one of learning to rank. We first formulate the problem of learning the mapping from the image features to the preference label as one of classification. In particular, we investigate the utilization of a multiple kernel learning algorithm based on group lasso to provide a solution. A simple but effective strategy to estimate perceptual image quality scores is then presented. Experiments show that the proposed BIQA method is highly effective and achieves a performance comparable with that of state-of-the-art BIQA algorithms. Moreover, the proposed method can be easily extended to new distortion categories.
Fei Gao 0006, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2013 Universal Blind Image Quality Assessment Metrics Via Natural Scene Statistics and Multiple Kernel Learning
abstract
Universal blind image quality assessment (IQA) metrics that can work for various distortions are of great importance for image processing systems, because neither ground truths are available nor the distortion types are aware all the time in practice. Existing state-of-the-art universal blind IQA algorithms are developed based on natural scene statistics (NSS). Although NSS-based metrics obtained promising performance, they have some limitations: 1) they use either the Gaussian scale mixture model or generalized Gaussian density to predict the nonGaussian marginal distribution of wavelet, Gabor, or discrete cosine transform coefficients. The prediction error makes the extracted features unable to reflect the change in nonGaussianity (NG) accurately. The existing algorithms use the joint statistical model and structural similarity to model the local dependency (LD). Although this LD essentially encodes the information redundancy in natural images, these models do not use information divergence to measure the LD. Although the exponential decay characteristic (EDC) represents the property of natural images that large/small wavelet coefficient magnitudes tend to be persistent across scales, which is highly correlated with image degradations, it has not been applied to the universal blind IQA metrics; and 2) all the universal blind IQA metrics use the same similarity measure for different features for learning the universal blind IQA metrics, though these features have different properties. To address the aforementioned problems, we propose to construct new universal blind quality indicators using all the three types of NSS, i.e., the NG, LD, and EDC, and incorporating the heterogeneous property of multiple kernel learning (MKL). By analyzing how different distortions affect these statistical properties, we present two universal blind quality assessment models, NSS global scheme and NSS two-step scheme. In the proposed metrics: 1) we exploit the NG of natural images using the original marginal distribution of wavelet coefficients; 2) we measure correlations between wavelet coefficients using mutual information defined in information theory; 3) we use features of EDC in universal blind image quality prediction directly; and 4) we introduce MKL to measure the similarity of different features using different kernels. Thorough experimental results on the Laboratory for Image and Video Engineering database II and the Tampere Image Database2008 demonstrate that both metrics are in remarkably high consistency with the human perception, and overwhelm representative universal blind algorithms as well as some standard full reference quality indexes for various types of distortions.
Xinbo Gao 0001, Fei Gao 0006, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2012 Local Structure Divergence Index for Image Quality Assessment
Fei Gao 0006, Dacheng Tao, Xuelong Li 0001, Xinbo Gao 0001, Lihuo He
ICONIP (5)1
2012 Color Fractal Structure Model for Reduced-Reference Colorful Image Quality Assessment
Lihuo He, Dongxue Wang, Xuelong Li 0001, Dacheng Tao, Xinbo Gao 0001, Fei Gao 0006
ICONIP (2)6
2010 An image quality assessment metric with no reference using hidden Markov tree model
abstract
No reference (NR) method is the most difficult issue of image quality assessment (IQA), which does not need the original image or its features as reference and only depends on the statistical law of the natural images. So, the NR-IQA is a high -level evaluation for image quality and simulates the complicated subjective process of human beings. This paper presents a NR-IQA metric based on Hidden Markov Tree (HMT) model. First, the HMT is utilized to model natural images, and the statistical properties of the model parameters are analyzed to mimic variation of image degradation. Then, by estimating the deviation degree of the parameters from the statistical law the distortion metric is constructed. Experimental results show that the proposed image quality assessment model is consistent well with the subjective evaluation results, and outperforms the existing models on difference distortions.
Fei Gao 0006, Xinbo Gao 0001, Wen Lu 0004, Dacheng Tao, Xuelong Li 0001
VCIP1