VLDB 2026 Research / reviewers in the wild / expert
Yuzhe Yang 0001
dblp:213/0962-1
· DBLP profile ↗
23ranked-venue papers
1as first author
21since 2021 · last 2025
0000-0001-9098-2105ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CLIP Brings Better Features to Visual Aesthetics LearnersabstractImage Aesthetics Assessment (IAA) is a challenging task due to its subjective nature and expensive manual annotations. Recent large-scale vision-language models, such as Contrastive Language-Image Pre-training (CLIP), have shown their promising representation capability for various downstream tasks. However, the application of CLIP to resource-constrained and low-data IAA tasks remains limited. While few attempts to leverage CLIP in IAA have mainly focused on carefully designed prompts, we extend beyond this by allowing models from different domains and with different model sizes to acquire knowledge from CLIP. To achieve this, we propose a unified and flexible two-phase CLIP-based Semi-supervised Knowledge Distillation (CSKD) paradigm, aiming to learn a lightweight IAA model while leveraging CLIP’s strong generalization capability. Specifically, CSKD employs a feature alignment strategy to facilitate the distillation of heterogeneous CLIP teacher and IAA student models, effectively transferring valuable features from pre-trained visual representations to two lightweight IAA models, respectively. To efficiently adapt to downstream IAA tasks in a low-data regime, the two strong visual aesthetics learners then conduct distillation with unlabeled examples for refining and transferring the task-specific knowledge collaboratively. Extensive experiments demonstrate that the proposed CSKD achieves state-of-the-art performance on multiple widely used IAA benchmarks. Furthermore, analysis of attention distance and entropy before and after feature alignment shows the effective transfer of CLIP’s feature representation to IAA models, which not only provides valuable guidance for the model initialization of IAA but also enhances the aesthetic feature representation of IAA models. Code will be made publicly available. Liwu Xu, Jinjin Xu, Yuzhe Yang 0001, Xilu Wang 0001, Yi-Jie Huang |
ICME | 3 |
| 2025 | Reproducibility Companion Paper: u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
Jinjin Xu, Xilu Wang 0001, Liwu Xu, Yuzhe Yang 0001, Xiang Li 0179, Fanyi Wang, Yanchun Xie, Yi-Jie Huang, Yunfan Hu |
ICMR | 4 |
| 2024 | u-LLaVA: Unifying Multi-Modal Tasks via Large Language ModelabstractRecent advancements in multi-modal large language models (MLLMs) have led to substantial improvements in visual understanding, primarily driven by sophisticated modality alignment strategies. However, predominant approaches prioritize global or regional comprehension, with less focus on fine-grained, pixel-level tasks. To address this gap, we introduce u-LLaVA, an innovative unifying multi-task framework that integrates pixel, regional, and global features to refine the perceptual faculties of MLLMs. We commence by leveraging an efficient modality alignment approach, harnessing both image and video datasets to bolster the model’s foundational understanding across diverse visual contexts. Subsequently, a joint instruction tuning method with task-specific projectors and decoders for end-to-end downstream training is presented. Furthermore, this work contributes a novel mask-based multi-task dataset comprising 277K samples, crafted to challenge and assess the fine-grained perception capabilities of MLLMs. The overall framework is simple, effective, and achieves state-of-the-art performance across multiple benchmarks. We make model, data, and code publicly accessible at https://github.com/OPPOMKLab/u-LLaVA. Jinjin Xu, Liwu Xu, Yuzhe Yang 0001, Xiang Li 0179, Fanyi Wang, Yanchun Xie, Yi-Jie Huang |
ECAI | 3 |
| 2024 | Emotion-aware hierarchical interaction network for multimodal image aesthetics assessment
Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001 |
Pattern Recognit. | 5 |
| 2024 | Coarse-to-Fine Image Aesthetics Assessment With Dynamic Attribute SelectionabstractImage aesthetics assessment (IAA) is an interesting but challenging task, owing to the ineffable nature of human sense of beauty. The study of IAA has evolved from simple binary classification to more complex score regression and distribution prediction. It is effortless for people to perform aesthetic binary classification,i.e., aesthetically pleasing or not. However, further judgment on the fine-level scalar aesthetic score is complex and typically determined by aesthetic attributes presented in the image, such as content, lighting and color. Motivated by the above facts, this paper presents a Coarse-to-fine image Aesthetics assessment model guided by Dynamic Attribute Selection, dubbed CADAS. The underlying idea is to simulate the process of human aesthetic perception by performing coarse-to-fine aesthetic reasoning. Specifically, a hierarchical AttributeNet is first pre-trained by imitating the staged mechanism of human aesthetic experience, producing the candidate aesthetic attributes. Then, an AestheticNet is introduced to perform the coarse-level binary classification, based on which a confidence-based attribute selection strategy is designed to dynamically pick out the dominant aesthetic attributes from the candidate ones. Finally, a self-attention-based FusionNet is designed to explore the interaction between dominant aesthetic attributes and aesthetic features, producing the fine-level aesthetic prediction. Extensive experiments demonstrate that the proposed model is superior to the state-of-the-arts. Furthermore, CADAS is also able to output the dominant aesthetic attributes in images, facilitating model explainability. Yipo Huang, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Guangming Shi |
IEEE Trans. Multim. | 5 |
| 2024 | Multi-Level Transitional Contrast Learning for Personalized Image Aesthetics AssessmentabstractPersonalized image aesthetics assessment (PIAA) is aimed at modeling the unique aesthetic preferences of individuals, based on which personalized aesthetic scores are predicted. People have different standards for image aesthetics, and accordingly, images rated at the same aesthetic level by different users explicitly reveal their aesthetic preferences. However, previous PIAA models treat each individual as an isolated optimization target, failing to take full advantage of the contrastive information among users. Further, although people's aesthetic preferences are unique, they still share some commonalities, meaning that PIAA models could be built on the basis of generic aesthetics. Motivated by the above facts, this article presents a Multi-level Transitional Contrast Learning (MTCL) framework for PIAA by transiting features from generic aesthetics to personalized aesthetics via contrastive learning. First, a generic image aesthetics assessment network is pre-trained to learn the common aesthetic features. Then, image sets rated to have the same aesthetic levels by different users are employed to learn the differentiated aesthetic features through multiple level-wise contrast learning based on the generic aesthetic features. Finally, a target user's PIAA model is built by integrating generic and differentiated aesthetic features. Extensive experiments on four benchmark PIAA databases demonstrate that the proposed MTCL model outperforms the state-of-the-arts. Zhichao Yang 0013, Leida Li, Yuzhe Yang 0001, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2023 | Mixed Sample Augmentation for Online DistillationabstractMixed Sample Regularization (MSR), such as MixUp or CutMix, is a powerful data augmentation strategy to generalize convolutional neural networks. Previous empirical analysis has illustrated an orthogonal performance gain between MSR and conventional offline Knowledge Distillation (KD). To be more specific, student networks can be enhanced with the involvement of MSR in the training stage of sequential distillation. Yet, the interplay between MSR and online knowledge distillation, where an ensemble of peer students learn mutually from each other, remains unexplored. To bridge the gap, we make the first attempt at incorporating CutMix into online distillation, where we empirically observe a significant improvement. Encouraged by this fact, we propose an even stronger MSR specifically for online distillation, named as CutnMix. Furthermore, a novel online distillation framework is designed upon CutnMix, to enhance the distillation with feature level mutual learning and a self-ensemble teacher. Comprehensive evaluations on CIFAR10 and CIFAR100 with six network architectures show that our approach can consistently outperform state-of-the-art distillation methods. Yiqing Shen 0003, Liwu Xu, Yuzhe Yang 0001, Yandong Guo |
ICASSP | 3 |
| 2023 | Attribute-assisted Multimodal Network for Image Aesthetics AssessmentabstractImage aesthetics assessment (IAA) is challenging due to its highly abstract nature. Nowadays, people tend to share images and comment them on social networks, which can provide rich information for judging image aesthetics. As a result, user comments of an image can be jointly utilized to learn better feature representations for IAA. Previous researches have shown that aesthetic attributes are crucial factors in determining image aesthetic quality and influencing people’s aesthetic perception. Accordingly, when commenting an image, people usually give descriptions from the perspective of aesthetic attributes. Inspired by this, this paper presents a new Attribute-Assisted Multimodal network (AAM-Net) for image aesthetics assessment. Specifically, we propose a cross-modal attribute interaction module to explore the related aesthetic attribute semantics shared by an image and the corresponding aesthetic comments. Then, a cross-modal gate unit is introduced to further refine significant attribute semantics interactively. Finally, informative aesthetic features can be obtained for predicting image aesthetic distributions. Experimental results on two public multimodal IAA databases demonstrate the superiority of the proposed model over the state-of-the-art methods. Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo |
ICME | 5 |
| 2023 | AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentabstractImage aesthetics assessment (IAA) aims at predicting the aesthetic quality of images. Recently, large pre-trained vision-language models, like CLIP, have shown impressive performances on various visual tasks. When it comes to IAA, a straightforward way is to finetune the CLIP image encoder using aesthetic images. However, this can only achieve limited success without considering the uniqueness of multimodal data in the aesthetics domain. People usually assess image aesthetics according to fine-grained visual attributes, e.g., color, light and composition. However, how to learn aesthetics-aware attributes from CLIP-based semantic space has not been addressed before. With this motivation, this paper presents a CLIP-based multi-attribute contrastive learning framework for IAA, dubbed AesCLIP. Specifically, AesCLIP consists of two major components, i.e., aesthetic attribute-based comment classification and attribute-aware learning. The former classifies the aesthetic comments into different attribute categories. Then the latter learns an aesthetic attribute-aware representation by contrastive learning, aiming to mitigate the domain shift from the general visual domain to the aesthetics domain. Extensive experiments have been done by using the pre-trained AesCLIP on four popular IAA databases, and the results demonstrate the advantage of AesCLIP over the state-of-the-arts. The source code will be public at https://github.com/OPPOMKLab/AesCLIP. Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong, Yuzhe Yang 0001, Liwu Xu, Guangming Shi |
ACM Multimedia | 6 |
| 2023 | Technical Quality-Assisted Image Aesthetics Quality Assessment
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Liwu Xu, Yuzhe Yang 0001 |
PRCV (11) | 6 |
| 2023 | Anchor-based knowledge embedding for image aesthetics assessment
Leida Li, Tianwu Zhi, Guangming Shi, Yuzhe Yang 0001, Liwu Xu, Yandong Guo |
Neurocomputing | 4 |
| 2023 | Theme-Aware Visual Attribute Reasoning for Image Aesthetics AssessmentabstractPeople usually assess image aesthetics according to visual attributes, e.g., interesting content, good lighting and vivid color, etc. Further, the perception of visual attributes depends on the image theme. Therefore, the inherent relationship between visual attributes and image theme is crucial for image aesthetics assessment (IAA), which has not been comprehensively investigated. With this motivation, this paper presents a new IAA model based on Theme-Aware Visual Attribute Reasoning (TAVAR). The underlying idea is to simulate the process of human perception in image aesthetics by performing bilevel reasoning. Specifically, a visual attribute analysis network and a theme understanding network are first pre-trained to extract aesthetic attribute features and theme features, respectively. Then, the first level Attribute-Theme Graph (ATG) is built to investigate the coupling relationship between visual attributes and image theme. Further, a flexible aesthetics network is introduced to extract general aesthetic features, based on which we built the second level Attribute-Aesthetics Graph (AAG) to mine the relationship between theme-aware visual attributes and aesthetic features, producing the final aesthetic prediction. Extensive experiments on four public IAA databases demonstrate the superiority of the proposed TAVAR model over the state-of-the-arts. Furthermore, TAVAR features better explainability due to the use of visual attributes. Leida Li, Yipo Huang, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Image Aesthetics Assessment With Attribute-Assisted Multimodal Memory NetworkabstractImage aesthetics assessment (IAA) has attracted growing interest in recent years but is still challenging due to its highly abstract nature. Nowadays, more and more people tend to comment images shared on the social networks, which can provide rich aesthetics-aware semantic information from different aspects. Therefore, user comments of an image can be exploited as supplementary information for enhancing aesthetic representation learning. Previous researches have demonstrated that aesthetic attributes make significant effect on image aesthetic quality and humans’ aesthetic perception. Typically, people are used to give comments on an image from the perspective of aesthetic attributes, based on which the aesthetic quality of images can be inferred. Motivated by this, this paper presents an Attribute-assisted Multimodal Memory Network (AMM-Net) for image aesthetics assessment, which utilizes aesthetic attributes to model the interactions between visual and textual modalities. Specifically, we design two memory networks to capture the attribute-aware information most related to the image and associated comments respectively. Further, with multiple memory hops, attribute semantics shared by the two modalities are refined and cross-modal interactions are enhanced progressively. Finally, more discriminative aesthetic representations can be obtained for IAA. The experimental results and comparisons on two public multimodal IAA datasets demonstrate the superiority of the proposed model over the state-of-the-art methods. The source code is available athttps://github.com/zhutong0219/AMM-Net. Leida Li, Tong Zhu 0003, Pengfei Chen 0003, Yuzhe Yang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Explainable and Generalizable Blind Image Quality Assessment via Semantic Attribute ReasoningabstractBlind image quality assessment (BIQA) that can directly evaluate image quality without perfect-quality reference has been a long-standing research topic. Although the existing BIQA models have achieved very encouraging performance, the lack of explainability and generalization ability limits their real-world applications to a great extent. People usually assess image quality according to semantic attributes, e.g., brightness, color, contrast, noise and sharpness. Furthermore, judgment on image quality is also impacted by the scene presented in the image. Therefore, the inherent relationship between semantic attributes and scenes is crucial for image quality assessment, which has rarely been explored yet. With this motivation, this paper presents a Semantic Attribute Reasoning based image QUality Evaluator (SARQUE). Specifically, we propose a two-stream network to predict semantic attributes and scene categories from distorted images. To investigate the inherent relationship between the semantic attributes and scene category, a semantic reasoning module is further proposed based on the graph convolution network (GCN), producing the final quality score. Extensive experiments conducted on five in-the-wild image quality databases demonstrate the superiority of the proposed SARQUE model over the state-of-the-arts. Furthermore, the proposed model features better explainability and generalization ability due to the use of semantic attributes. Yipo Huang, Leida Li, Yuzhe Yang 0001, Yandong Guo |
IEEE Trans. Multim. | 3 |
| 2023 | Knowledge-Guided Blind Image Quality Assessment With Few Training SamplesabstractBlind image quality assessment (BIQA) for in-the-wild images has achieved great progress by training advanced deep neural networks. However, the current BIQA models are suffering the generalization challenge, meaning that a well-trained BIQA model is still very limited in evaluating images with different distributions. Deep BIQA models are data-intensive, but the annotation of image quality labels is extremely expensive. To design a generalizable BIQA model with few training samples is highly desired. Motivated by the above fact, this paper presents a knowledge-guided BIQA (KG-IQA) framework by integrating domain knowledge from the human visual system (HVS) and natural scene statistics (NSS). Specifically, the quality-aware HVS and NSS features are first extracted as prior knowledge. Then, we embed the two types of knowledge into the conventional deep neural network by learning to predict the HVS and NSS features, producing the knowledge-enhanced quality features, based on which the final image quality score is obtained. We conduct extensive experiments and comparisons on five authentically distorted IQA datasets. The experimental results demonstrate that the introduction of knowledge greatly reduces the requirement on the amount of training images, and the proposed KG-IQA model achieves superior performance in terms of both prediction accuracy and generalization ability. Tianshu Song, Leida Li, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi |
IEEE Trans. Multim. | 4 |
| 2022 | Self-Distillation from the Last Mini-Batch for Consistency RegularizationabstractKnowledge distillation (KD) shows a bright promise as a powerful regularization strategy to boost generalization ability by leveraging learned sample-level soft targets. Yet, employing a complex pre-trained teacher network or an ensemble of peer students in existing KD is both timeconsuming and computationally costly. Various self KD methods have been proposed to achieve higher distillation efficiency. However, they either require extra network architecture modification or are difficult to parallelize. To cope with these challenges, we propose an efficient and reliable self-distillation framework, named Self-Distillation from Last Mini-Batch (DLB). Specifically, we rearrange the sequential sampling by constraining half of each mini-batch coinciding with the previous iteration. Meanwhile, the rest half will coincide with the upcoming iteration. Afterwards, the former half mini-batch distills on-the-fly soft targets generated in the previous iteration. Our proposed mechanism guides the training stability and consistency, resulting in robustness to label noise. Moreover, our method is easy to implement, without taking up extra run-time memory or requiring model structure modification. Experimental results on three classification benchmarks illustrate that our approach can consistently outperform state-of-the-art self-distillation approaches with different network architectures. Additionally, our method shows strong compatibility with augmentation strategies by gaining additional performance improvement. The code is available at https://github.com/Meta-knowledge-Lab/DLB. Yiqing Shen 0003, Liwu Xu, Yuzhe Yang 0001, Yandong Guo |
CVPR | 3 |
| 2022 | Personalized Image Aesthetics Assessment with Rich AttributesabstractPersonalized image aesthetics assessment (PIAA) is challenging due to its highly subjective nature. People's aesthetic tastes depend on diversified factors, including image characteristics and subject characters. The existing PIAA databases are limited in terms of annotation diversity, especially the subject aspect, which can no longer meet the increasing demands of PIAA research. To solve the dilemma, we conduct so far, the most comprehensive subjective study of personalized image aesthetics and introduce a new Personalized image Aesthetics database with Rich Attributes (PARA), which consists of 31,220 images with annotations by 438 subjects. PARA features wealthy annotations, including 9 image-oriented objective attributes and 4 human-oriented subjective attributes. In addition, desensitized subject information, such as personality traits, is also provided to support study of PIAA and user portraits. A comprehensive analysis of the annotation data is provided and statistic study indicates that the aesthetic preferences can be mirrored by proposed subjective attributes. We also propose a conditional PIAA model by utilizing subject information as conditional prior. Experimental results indicate that the conditional PIAA model can outperform the control group, which is also the first attempt to demonstrate how image aesthetics and subject characters interact to produce the intricate personalized tastes on image aesthetics. We believe the database and the associated analysis would be useful for conducting next-generation PIAA study. The project page of PARA can be found at: https://cv-datasets.institutecv.com/#/data-sets. Yuzhe Yang 0001, Liwu Xu, Leida Li, Nan Qie, Yandong Guo |
CVPR | 1 |
| 2022 | Psychology Inspired Model for Hierarchical Image Aesthetic Attribute PredictionabstractDeep neural network has proved its effectiveness in image aesthetic quality assessment (IAQA), but still lacks reasonable interpretability. Aesthetic attributes provide rich intermediate-level information for understanding the underlying principles of image aesthetics, but has not been fully investigated. Psychological studies have shown that aesthetic experience involves hierarchical stages, i.e., human process image aesthetics following a staged information processing mechanism. Motivated by this, this paper presents a Hierarchical Image Aesthetic Attribute (HIAA) prediction model, aiming to imitate the staged mechanism of human aesthetic experience. Image aesthetic attributes are first divided into several hierarchical groups. Then, hierarchical features are extracted from the cascaded layers of the deep neural network to predict the aesthetic attributes in a group-wise manner. The overall image aesthetic score is also predicted by aggregating the hierarchical features. Experimental results demonstrate that the proposed HIAA model outperforms the state-of-the-arts in terms of both aesthetic attribute prediction and aesthetic score regression. Leida Li, Jiachen Duan, Yuzhe Yang 0001, Liwu Xu, Yandong Guo |
ICME | 3 |
| 2022 | Transductive Aesthetic Preference Propagation for Personalized Image Aesthetics AssessmentabstractPersonalized image aesthetics assessment (PIAA) aims at capturing individual aesthetic preference. Fine-tuning on personalized data has been proven to be effective in PIAA task. However, a fixed fine-tuning strategy may cause under/over-fitting on limited personal data and it also brings additional training cost. To alleviate these issues, we employ a meta learning-based Transductive Aesthetic Preference Propagation (TAPP-PIAA) algorithm under regression manner to substitute the fine-tuning strategy. Specifically, each user's data is regarded as a meta-task and spilt into support and query set. Then, we extract deep aesthetic features with a pre-trained generic image aesthetic assessment (GIAA) model. Next, we treat image features as graph nodes and their similarities as edge weights to construct an undirected nearest neighbor graph for inference. Instead of fine-tuning on support set, TAPP-PIAA propagates aesthetic preference from support to query set with a predefined propagation formula. Finally, to learn a generalizable aesthetic representation for various users, we optimize our TAPP-PIAA across different users with meta-learning framework. Experimental results indicate that our TAPP-PIAA can surpass the state-of-the-art methods on benchmark databases. Yuzhe Yang 0001, Huaxiong Li, Haoxing Chen, Liwu Xu, Leida Li, Yandong Guo |
ACM Multimedia | 2 |
| 2022 | Learning image aesthetic subjectivity from attribute-aware relational reasoning network
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Guangcheng Wang, Yuzhe Yang 0001 |
Pattern Recognit. Lett. | 5 |
| 2022 | Hierarchical discrepancy learning for image restoration quality assessment
Bo Hu 0008, Shuaijian Wang, Leida Li, Jiaxu Leng, Yuzhe Yang 0001, Xinbo Gao 0001 |
Signal Process. | 5 |
| 2020 | Blind Realistic Blur Assessment Based on Discrepancy LearningabstractBlur is one of the most common distortions that degrade natural images. This stimulates the blossom of sharpness assessment metrics. Existing sharpness metrics possess good performance for evaluating simulated blur, but are limited for the more common realistic blur that are introduced during image capture and processing in real life. To this end, we propose an effective Realistic Blur Assessment method (RBA) based on discrepancy learning. First, motivated by the fact that the distortion-free reference images are usually unavailable in practice, but the Human Visual System (HVS) can still accurately perceive image sharpness by quantifying the perceptual discrepancy between the distorted image and the hallucinated reference image in mind, we propose to train a discrepancy generation model to automatically generate the discrepancy map from the distorted image analogous to the HVS. This is achieved by using a deep neural network with rich training images. With the discrepancy map, two sharpness-aware features, i.e. sparse representation based entropy of primitive and content-guided variation of power, are then extracted to severally quantify spatial visual information amount and spectral power. Finally, the two features are integrated to produce the overall sharpness score. Extensive experiments demonstrate the superiority of the proposed method over the state-of-the-arts. Leida Li, Yu Zhou 0009, Ke Gu 0001, Yuzhe Yang 0001, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Naturalness Preserved Image Aesthetic Enhancement with Perceptual Encoder ConstraintabstractTypical supervised image enhancement pipeline is to minimize the distance between the enhanced image and the reference one. Pixel-wise and perceptual-wise loss functions could help to improve the general image quality, however are not very efficient in improving the image aesthetic quality. In this paper, we propose a novel Residual connected Dilated U-Net (RDU-Net) for improving the image aesthetic quality. By using different dilation rates, the RDU-Net can extract multiple receptive-field features and merge the maximum information from local to global, which are highly desired in image enhancement. Also, we propose an encoder constraint perceptual loss, which could teach the enhancement network to dig out the latent aesthetic factors and make the enhanced image more natural and aesthetically appealing. The proposed approach can alleviate the over-enhancement phenomenons. The experimental results show that the proposed perceptual loss function could give a steady back propagation and the proposed method outperforms the state-of-the-arts. Leida Li, Yuzhe Yang 0001, Hancheng Zhu |
ICMR | 2 |