Pengfei Chen 0003

dblp:13/3788-3 · DBLP profile ↗
← Back
43ranked-venue papers
14as first author
39since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 11 first-author · 33 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning
abstract
Livestreaming has become increasingly prevalent in modern visual communication, where automatic camera quality tuning is essential for delivering superior user Quality of Experience (QoE). Such tuning requires accurate blind image quality assessment (BIQA) to guide parameter optimization decisions. Unfortunately, the existing BIQA models typically only predict an overall coarse-grained quality score, which cannot provide fine-grained perceptual guidance for precise camera parameter tuning. To bridge this gap, we first establish FGLive-10K, a comprehensive fine-grained BIQA database containing 10,185 high-resolution images captured under varying camera parameter configurations across diverse livestreaming scenarios. The dataset features 50,925 multi-attribute quality annotations and 19,234 fine-grained pairwise preference annotations. Based on FGLive-10K, we further develop TuningIQA, a fine-grained BIQA metric for livestreaming camera tuning, which integrates human-aware feature extraction and graph-based camera parameter fusion. Extensive experiments and comparisons demonstrate that TuningIQA significantly outperforms state-of-the-art BIQA methods in both score regression and fine-grained quality ranking, achieving superior performance when deployed for livestreaming camera tuning.
Xiangfei Sheng, Zhichao Duan 0002, Xiaofeng Pan, Yipo Huang, Zhichao Yang 0013, Pengfei Chen 0003, Leida Li
AAAI6
2026 Fine-grained Image Quality Assessment for Perceptual Image Restoration
abstract
Recent years have witnessed remarkable achievements in perceptual image restoration (IR), creating an urgent demand for accurate image quality assessment (IQA), which is essential for both performance comparison and algorithm optimization. Unfortunately, the existing IQA metrics exhibit inherent weakness for IR task, particularly when distinguishing fine-grained quality differences among restored images. To address this dilemma, we contribute the first-of-its-kind fine-grained image quality assessment dataset for image restoration, termed FGRestore, comprising 18,408 restored images across six common IR tasks. Beyond conventional scalar quality scores, FGRestore was also annotated with 30,886 fine-grained pairwise preferences. Based on FGRestore, a comprehensive benchmark was conducted on the existing IQA metrics, which reveal significant inconsistencies between score-based IQA evaluations and the fine-grained restoration quality. Motivated by these findings, we further propose FGResQ, a new IQA model specifically designed for image restoration, which features both coarse-grained score regression and fine-grained quality ranking. Extensive experiments and comparisons demonstrate that FGResQ significantly outperforms state-of-the-art IQA metrics.
Xiangfei Sheng, Xiaofeng Pan, Zhichao Yang 0013, Pengfei Chen 0003, Leida Li
AAAI4
2026 LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations
abstract
The increasing popularity of long Text-to-Image (T2I) generation has created an urgent need for automatic and interpretable models that can evaluate the image-text alignment in long prompt scenarios. However, the existing T2I alignment benchmarks predominantly focus on short prompt scenarios and only provide MOS or Likert scale annotations. This inherent limitation hinders the development of long T2I evaluators, particularly in terms of the interpretability of alignment. In this study, we contribute LongT2IBench, which comprises 14K long text-image pairs accompanied by graph-structured human annotations. Given the detail-intensive nature of long prompts, we first design a Generate-Refine-Qualify annotation protocol to convert them into textual graph structures that encompass entities, attributes, and relations. Through this transformation, fine-grained alignment annotations are achieved based on these granular elements. Finally, the graph-structed annotations are converted into alignment scores and interpretations to facilitate the design of T2I evaluation models. Based on LongT2IBench, we further propose LongT2IExpert, a LongT2I evaluator that enables multi-modal large language models (MLLMs) to provide both quantitative scores and structured interpretations through an instruction-tuning process with Hierarchical Alignment Chain-of-Thought (CoT). Extensive experiments and comparisons demonstrate the superiority of the proposed LongT2IExpert in alignment evaluation and interpretation.
Zhichao Yang 0013, Tianjiao Gu, Jianjie Wang, Feiyu Lin, Xiangfei Sheng, Pengfei Chen 0003, Leida Li
AAAI6
2026 HumanCrop-Thinker: An inference-driven framework with explicit thinking for explainable human-centric image cropping
Yipo Huang, Pengfei Chen 0003, Leida Li
Expert Syst. Appl.3
2026 Blind Omnidirectional Image Quality Assessment: Embracing the Magic Power of Multimodal Large Language Models
Jiebin Yan, Junjie Chen 0008, Pengfei Chen 0003, Xuelin Liu, Ziwen Tan, Yuming Fang 0001
Int. J. Comput. Vis.4
2026 Learning Scene-Invariant Distribution for Generalizable Blind Image Quality Assessment
abstract
The inherent diversity of visual scenes poses a fundamental challenge in blind image quality assessment (BIQA), which has become a major obstacle to the model generalization. In this study, we found that human annotations for images with different visual scenes exhibit distinct quality distribution discrepancies. The existing BIQA models tend to overfit to such diversified distributions, which in turn leads to compromised model generalizability, especially when dealing with unseen scenes in the real-world scenario. Motivated by the above facts, this paper presents a generalizable BIQA model by learning Scene-INvariant Distribution, named SIND. Specifically, we propose a distribution alignment framework to alleviate the distribution discrepancy for quality regression models, which is achieved by automatically scaling and shifting the cross-scene distributions into a unified distribution. Then, the aligned unified distribution is leveraged to supervise the model training, achieving scene-invariant and quality-aware feature representation. In addition, a token-complementary patch reasoning network is designed to extract comprehensive quality-aware features from both the image overview and detail, achieving more accurate quality prediction. Extensive experiments for both image technical- and aesthetic-quality assessment tasks show the superiority of the proposed SIND model over the state-of-the-arts. Moreover, the proposed framework is model-agnostic and can enhance model generalizability without incurring extra inference costs. The proposed method won the championship in the NTIRE 2024 Portrait Quality Assessment Challenge. Codes will be available at https://github.com/ZachL1/SIND.
Yipo Huang, Zhichao Duan 0002, Pengfei Chen 0003, Leida Li, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.3
2026 Viewport-Unaware Full-Reference Omnidirectional Image Quality Assessment With Inter-Patch and Sequence Similarity
abstract
Full-reference (FR) image quality assessment (IQA) (FR-IQA) has been extensively explored in the past two decades and is one of the most basic and hot topics in the image processing community, due to its indispensable role in quantitatively describing image quality degradation and guiding algorithm and system optimization. However, FR omnidirectional image quality assessment (OIQA) (FR-OIQA) has achieved less success, due to the natural gap between 2D images and omnidirectional images (OIs). To this end, we present a novel FR-OIQA model with Inter-Patch and Sequence Similarity (IPSS). Specifically, to avoid the extra computational load of viewport generation/prediction methods, IPSS processes OIs in aviewport-unawaremanner,i.e., directly extracting a patch sequence from an OI in the format of Equirectangular Projection (ERP) with retaining regions of interest. Furthermore, since the patches from ERP image contain inborn geometry deformation, thedeformation-awareconvolution is plugged into feature extraction and used to distill quality-aware features from theintrinsic pseudo-degradation, which are then utilized to measure inter-patch similarity. Finally, a distortion-aware interaction module is used to aggregate patch-wise quality-aware features, whose output is used to calculate patch-sequence similarity,i.e., the global quality of OI. Through comprehensive experiments on a large-scale OIQA database, we demonstrate the superiority of the proposed IPSS and the effectiveness of each module.
Jiebin Yan, Junjie Chen 0008, Pengfei Chen 0003, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 RAM-VQA: Restoration Assisted Multi-Modality Video Quality Assessment
abstract
Video Quality Assessment (VQA) strives to computationally emulate human perceptual judgments and has garnered significant attention given its widespread applicability. However, existing methodologies face two primary impediments: (1) limited proficiency in evaluating samples at quality extremes (e.g., severely degraded or near-perfect videos), and (2) insufficient sensitivity to nuanced quality variations arising from a misalignment with human perceptual mechanisms. Although vision-language models offer promising semantic understanding, their reliance on visual encoders pre-trained for high-level tasks often compromises their sensitivity to low-level distortions. To surmount these challenges, we propose the Restoration-Assisted Multi-modality VQA (RAM-VQA) framework. Uniquely, our approach leverages video restoration as a proxy to explicitly model distortion-sensitive features. The framework operates through two synergistic stages: a prompt learning stage that constructs a quality-aware textual space using triple-level references (degraded, restored, and pristine) derived from the restoration process, and a dual-branch evaluation stage that integrates semantic cues with technical quality indicators via spatio-temporal differential analysis. Extensive experiments demonstrate that RAM-VQA achieves state-of-the-art performance across diverse benchmarks, exhibiting superior capability in handling extreme-quality content while ensuring robust generalization.
Pengfei Chen 0003, Jiebin Yan, Rajiv Soundararajan, Giuseppe Valenzise, Leida Li
IEEE Trans. Image Process.1
2026 Exploring Cross-Modal Mutual Prompt Learning for Video Quality Assessment
abstract
Enhancing video quality assessment (VQA) through semantic information integration is a critical research focus. Recent research has employed the Contrastive Language-Image Pre-training (CLIP) model as a foundation to improve semantic perception. However, the image-text alignment inherent in these pre-trained Vision-Language (VL) models frequently results in suboptimal VQA performance. While prompt engineering has recently targeted the language component to address this alignment issue, the unique insights resided in visual analysis is still overlooked for further advancing VQA tasks. Additionally, seeking a trade-off between quality separability and domain invariance in VQA remains largely unresolved within the VL paradigm. In this paper, we introduce a novel cross-modal prompt-based approach to tackle these challenges. Specifically, we propose learnable prompts within the vision branch to foster synergy between visual and language modalities through a language-to-vision coupling function. The multi-view backbone is then carefully crafted with content enhancement and distortion-aware temporal modulation to ensure quality separability. The language prompts, derived from visual representations, are further supported by adaptive weighting mechanisms to optimize the balance between quality separability and domain invariance. Experimental results demonstrate the effectiveness of our proposed method over leading VQA models, showing significant improvements in generalization across diverse datasets. The source code for this work is publicly available athttps://github.com/cpf0079/CM2PL.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Jiebin Yan, Vinit Jakhetiya, Aladine Chetouani
IEEE Trans. Multim.1
2026 AesPrompt: Zero-Shot Image Aesthetics Assessment With Multi-Granularity Aesthetic Prompt Learning
abstract
Recent years have witnessed increasing interest towards image aesthetics assessment (IAA), which predicts the aesthetic appeal of images by simulating human perception. The state-of-the-art IAA methods, despite their significant advancements, typically rely heavily on time-consuming and labor-intensive human annotation of aesthetic scores. Furthermore, they are subject to the generalization challenge, which is highly desired in real-world applications. Motivated by this, zero-shot image aesthetics assessment (ZIAA) is investigated to achieve robust model generalization without relying on manual aesthetic annotations, which remains largely underexplored. Specifically, a novel aesthetic prompt learning framework for ZIAA, dubbed AesPrompt, is presented in this paper. The key insight of AesPrompt is to emulate the human aesthetic perception process for learning aesthetic-oriented prompts in a multi-granularity manner. First, we first develop a new pseudo aesthetic distribution generation paradigm based on multi-LLM ensemble. Then, external knowledge of multi-granularity prompts encompassing image themes, emotions, and aesthetics is acquired. Through learning the multi-granularity aesthetic-oriented prompts, the proposed method achieves better generalization and interpretability. Extensive experiments on five IAA benchmarks demonstrate that AesPrompt consistently outperforms the state-of-the-art ZIAA methods across diverse-sourced images, covering natural images, artistic images, and artificial intelligence-generated images.
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Giuseppe Valenzise
IEEE Trans. Multim.3
2025 Text-to-Image Diffusion Models are AI-Generated Image Quality Scorers
abstract
Despite the remarkable progress in text-to-image generation models, AI-Generated Images (AGIs) still face significant challenges, such as poor perception quality and text-image misalignment. Consequently, AI-Generated Image Quality Assessment (AGIQA), which aims to assess capabilities of generative models, has gained increasing attention. Diffusion models, pre-trained on large-scale text-image generative tasks, inherently encode rich prior knowledge regarding image quality and text-image alignment, which is crucial for AGIQA task. Inspired by this, unlike most existing CLIP-based methods, we present an initial exploration of diffusion model-based AGIQA, termed Diff-AGIQA. Specifically, to effectively harness pre-trained diffusion models, we introduce a simple yet effective strategy focusing on two key aspects: 1) feature selection, identifying the discriminative features within diffusion models; and 2) visual prompting, to extract more task-specific features while keeping the diffusion model frozen. Extensive experiments on AGIQA-1K, AGIQA-3K, AGIQA-20K, and RichHF-18K datasets demonstrate the potential of applying diffusion models to AGIQA task. Code is publicly available https://github.com/sxfly99/Diff-AGIQA.
Xiangfei Sheng, Weidong Zou, Pengfei Chen 0003, Leida Li
ICME3
2025 InstructCrop: Teaching Multimodal Large Language Models to Crop Aesthetic Images
abstract
Aesthetic Image Cropping (AIC) aims to improve the visual appeal of images by removing redundant content while preserving attractive elements. Despite the encouraging progresses achieved in data-driven approaches, most existing models struggle to understand user intentions, particularly for diversified scenes with multiple subjects. Moreover, they can only provide cropping results without explanations, which further restricts their usability in real-world applications. Motivated by the above facts, we introduce InstructCrop : a multimodal large language model (MLLM)-based AIC framework, which can understand user instructions and provide explanatory reasons for cropping results. Specifically, we first build a multimodal Image Cropping Instruction Tuning (ICIT) dataset through a cost-effective paradigm by generating high-quality instruction tuning data based on the existing cropping datasets. Then, we embed dynamic domain knowledge into the cropping model by integrating cropping-aware experts of aesthetic assessment and composition classification. Finally, we adapt MLLMs to generate the cropping results and corresponding explanations. Quantitative and qualitative experiments on three benchmark datasets demonstrate that InstructCrop enables effective and interpretable image cropping, which aligns better with user intentions. Data and code are available at https://github.com/sxfly99/InstructCrop.
Xiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen 0003, Tong Zhu 0003, Leida Li
ACM Multimedia4
2025 Multi-Modality Multi-Attribute Contrastive Pre-Training for Image Aesthetics Computing
abstract
In the Image Aesthetics Computing (IAC) field, most prior methods leveraged the off-the-shelf backbones pre-trained on the large-scale ImageNet database. While these pre-trained backbones have achieved notable success, they often overemphasize object-level semantics and fail to capture the high-level concepts of image aesthetics, which may only achieve suboptimal performances. To tackle this long-neglected problem, we propose a multi-modality multi-attribute contrastive pre-training framework, targeting at constructing an alternative to ImageNet-based pre-training for IAC. Specifically, the proposed framework consists of two main aspects. 1) We build a multi-attribute image description database with human feedback, leveraging the competent image understanding capability of the multi-modality large language model to generate rich aesthetic descriptions. 2) To better adapt models to aesthetic computing tasks, we integrate the image-based visual features with the attribute-based text features, and map the integrated features into different embedding spaces, based on which the multi-attribute contrastive learning is proposed for obtaining more comprehensive aesthetic representation. To alleviate the distribution shift encountered when transitioning from the general visual domain to the aesthetic domain, we further propose a semantic affinity loss to restrain the content information and enhance model generalization. Extensive experiments demonstrate that the proposed framework sets new state-of-the-arts for IAC tasks.
Yipo Huang, Leida Li, Pengfei Chen 0003, Haoning Wu 0001, Weisi Lin, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Towards Explainable Image Aesthetics Assessment With Attribute-Oriented Critiques Generation
abstract
Compared with the unimodal image aesthetics assessment (IAA), multimodal IAA has demonstrated superior performance. This indicates that the critiques could provide rich aesthetics-aware semantic information, which also enhance the explainability of IAA models. However, images are not always accompanied with critiques in real-world situation, rendering multimodal IAA inapplicable in most cases. Therefore, it would be interesting to investigate whether we can generate aesthetic critiques to facilitate image aesthetic representation learning and enhance model explainability. Motivated by these facts, this paper presents an attribute-oriented Critiques Generation framework for explainable IAA, dubbed CG-IAA, which consists of three major components, i.e., Vision-Language Aesthetic Pretraining (VLAP), Multi-Attribute Experts Learning (MAEL) and Multimodal Aesthetics Prediction (MAP). Specifically, the vanilla CLIP is first finetuned on a multimodal IAA database. Considering that the aesthetic critiques typically consist of multiple attributes, a new multimodal IAA database which contains over 1 million critiques with up to four aesthetic attributes is constructed with the language model-based knowledge transfer. Then, CLIP-based multi-attribute experts are trained based on this database. Finally, the pretrained experts are utilized to generate aesthetic critiques for assisting unimodal image aesthetics prediction. Extensive experiments have been done on four popular IAA databases, and the results demonstrate the advantage of CG-IAA over the state-of-the-arts. Furthermore, CG-IAA features better explainability and generalization with the assistance of generated critiques. The source code is available athttps://github.com/sxfly99/CG-IAA.
Leida Li, Xiangfei Sheng, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong
IEEE Trans. Circuits Syst. Video Technol.3
2024 Semantics-Aware Image Aesthetics Assessment using Tag Matching and Contrastive Ranking
abstract
The perception of image aesthetics is built upon the understanding of semantic content. However, how to evaluate the aesthetic quality of images with diversified semantic backgrounds remains challenging in image aesthetics assessment (IAA). To address the dilemma, this paper presents a semantics-aware image aesthetics assessment approach, which first analyzes the semantic content of images and then models the aesthetic distinctions among images from two perspectives, i.e., aesthetic attribute and aesthetic level. Concretely, we propose two strategies, dubbed tag matching and contrastive ranking, to extract knowledge pertaining to image aesthetics. The tag matching identifies the semantic category and the dominant aesthetic attributes based on predefined tag libraries. The contrastive ranking is designed to uncover the comparative relationships among images with different aesthetic levels but similar semantic backgrounds. In the process of contrastive ranking, the impact of long-tailed distribution of aesthetic data is also considered by balanced sampling and traversal contrastive learning. Extensive experiments and comparisons on three benchmark IAA databases demonstrate the superior performance of the proposed model in terms of both prediction accuracy and alleviating long-tailed effect. The code will be public at https://github.com/yzc-ippl/TMCR **REMOVE 2nd URL**://github.com/yzc-ippl/TMCR.
Zhichao Yang 0013, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong
ACM Multimedia3
2024 AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception
abstract
The highly abstract nature of image aesthetics perception (IAP) poses a significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma, resulting in MLLMs falling short of aesthetics perception capabilities. To address the above challenge, we first introduce a comprehensively annotated Aesthetic Multi-Modality Instruction Tuning (AesMMIT) dataset, which serves as the footstone for building multi-modality aesthetics foundation models. Specifically, to align MLLMs with human aesthetics perception, we construct a corpus-rich aesthetic critique database with 21,904 diverse-sourced images and 88K human natural language feedbacks, which are collected via progressive questions, ranging from coarse-grained aesthetic grades to fine-grained aesthetic descriptions. To ensure that MLLMs can handle diverse queries, we further prompt GPT to refine the aesthetic critiques and assemble the large-scale aesthetic instruction tuning dataset, i.e. AesMMIT, which consists of 409K multi-typed instructions to activate stronger aesthetic capabilities. Based on the AesMMIT database, we fine-tune the open-sourced general foundation models, achieving multi-modality Aesthetic Expert models, dubbed AesExpert. Extensive experiments demonstrate that the proposed AesExpert models deliver significantly better aesthetic perception performances than the state-of-the-art MLLMs, including the most advanced GPT-4V and Gemini-Pro-Vision. Project Page: https://yipoh.github.io/aes-expert/.
Yipo Huang, Xiangfei Sheng, Zhichao Yang 0013, Zhichao Duan 0002, Pengfei Chen 0003, Leida Li, Weisi Lin, Guangming Shi
ACM Multimedia6
2024 Aesthetic image cropping meets VLP: Enhancing good while reducing bad
Leida Li, Pengfei Chen 0003
J. Vis. Commun. Image Represent.3
2024 Emotion-aware hierarchical interaction network for multimodal image aesthetics assessment
Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001
Pattern Recognit.3
2024 Active Learning-Based Sample Selection for Label-Efficient Blind Image Quality Assessment
abstract
Despite the considerable effort devoted to high-generalizable blind image quality assessment (BIQA), the generalization performance of the state-of-the-art metrics remains limited when facing new visual scenes. A straightforward way to address the dilemma is labeling a great number of images from the new scene and subsequently training a new model, which is quite labor-intensive and cost-expensive. Hence, there is an urgent need to mitigate the dependency on labeled samples by designing a data-efficient BIQA algorithm. Motivated by the above facts, this paper presents an Active Learning-based IQA (AL-IQA) framework, which reduces the requirement for training samples by selecting representative images from two perspectives, including distortion and content. Specifically, in terms of distortion, we design distortion prompts and adopt Contrastive Language-Image Pre-Training (CLIP) to predict image distortion in a zero-shot manner. Then, we employ curriculum learning-inspired strategy to select samples with gradually increasing difficulty (measured by prediction uncertainty of CLIP), in order to facilitate model training. Meantime, in terms of content, we adopt distribution matching-based dataset distillation to distill unlabeled images into several high-density informative synthetic images. Then, feature distances between unlabeled images and distilled images are compared to identify images with the most representative content. Finally, Borda count is adopted to capture a consensus of both distortion and content through weighted counting, and prompt tuning is utilized for adapting the model to the IQA task. Extensive experiments are conducted on five IQA datasets, and the results demonstrate that the proposed AL-IQA not only effectively reduces the number of training samples but also achieves state-of-the-art prediction accuracy and generalization performance. The source code is available athttps://github.com/esnthere/AL-IQA.
Tianshu Song, Leida Li, Deqiang Cheng 0001, Pengfei Chen 0003, Jinjian Wu
IEEE Trans. Circuits Syst. Video Technol.4
2024 Coarse-to-Fine Image Aesthetics Assessment With Dynamic Attribute Selection
abstract
Image aesthetics assessment (IAA) is an interesting but challenging task, owing to the ineffable nature of human sense of beauty. The study of IAA has evolved from simple binary classification to more complex score regression and distribution prediction. It is effortless for people to perform aesthetic binary classification,i.e., aesthetically pleasing or not. However, further judgment on the fine-level scalar aesthetic score is complex and typically determined by aesthetic attributes presented in the image, such as content, lighting and color. Motivated by the above facts, this paper presents a Coarse-to-fine image Aesthetics assessment model guided by Dynamic Attribute Selection, dubbed CADAS. The underlying idea is to simulate the process of human aesthetic perception by performing coarse-to-fine aesthetic reasoning. Specifically, a hierarchical AttributeNet is first pre-trained by imitating the staged mechanism of human aesthetic experience, producing the candidate aesthetic attributes. Then, an AestheticNet is introduced to perform the coarse-level binary classification, based on which a confidence-based attribute selection strategy is designed to dynamically pick out the dominant aesthetic attributes from the candidate ones. Finally, a self-attention-based FusionNet is designed to explore the interaction between dominant aesthetic attributes and aesthetic features, producing the fine-level aesthetic prediction. Extensive experiments demonstrate that the proposed model is superior to the state-of-the-arts. Furthermore, CADAS is also able to output the dominant aesthetic attributes in images, facilitating model explainability.
Yipo Huang, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Guangming Shi
IEEE Trans. Multim.3
2023 Attribute-assisted Multimodal Network for Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) is challenging due to its highly abstract nature. Nowadays, people tend to share images and comment them on social networks, which can provide rich information for judging image aesthetics. As a result, user comments of an image can be jointly utilized to learn better feature representations for IAA. Previous researches have shown that aesthetic attributes are crucial factors in determining image aesthetic quality and influencing people’s aesthetic perception. Accordingly, when commenting an image, people usually give descriptions from the perspective of aesthetic attributes. Inspired by this, this paper presents a new Attribute-Assisted Multimodal network (AAM-Net) for image aesthetics assessment. Specifically, we propose a cross-modal attribute interaction module to explore the related aesthetic attribute semantics shared by an image and the corresponding aesthetic comments. Then, a cross-modal gate unit is introduced to further refine significant attribute semantics interactively. Finally, informative aesthetic features can be obtained for predicting image aesthetic distributions. Experimental results on two public multimodal IAA databases demonstrate the superiority of the proposed model over the state-of-the-art methods.
Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo
ICME3
2023 AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) aims at predicting the aesthetic quality of images. Recently, large pre-trained vision-language models, like CLIP, have shown impressive performances on various visual tasks. When it comes to IAA, a straightforward way is to finetune the CLIP image encoder using aesthetic images. However, this can only achieve limited success without considering the uniqueness of multimodal data in the aesthetics domain. People usually assess image aesthetics according to fine-grained visual attributes, e.g., color, light and composition. However, how to learn aesthetics-aware attributes from CLIP-based semantic space has not been addressed before. With this motivation, this paper presents a CLIP-based multi-attribute contrastive learning framework for IAA, dubbed AesCLIP. Specifically, AesCLIP consists of two major components, i.e., aesthetic attribute-based comment classification and attribute-aware learning. The former classifies the aesthetic comments into different attribute categories. Then the latter learns an aesthetic attribute-aware representation by contrastive learning, aiming to mitigate the domain shift from the general visual domain to the aesthetics domain. Extensive experiments have been done by using the pre-trained AesCLIP on four popular IAA databases, and the results demonstrate the advantage of AesCLIP over the state-of-the-arts. The source code will be public at https://github.com/OPPOMKLab/AesCLIP.
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong, Yuzhe Yang 0001, Liwu Xu, Guangming Shi
ACM Multimedia3
2023 Personalized Image Aesthetics Assessment with Attribute-guided Fine-grained Feature Representation
abstract
Personalized image aesthetics assessment (PIAA) has gained increasing attention from researchers due to its ability to measure individual users' specific aesthetic experiences. However, most existing PIAA methods rely on holistic features or simplistic coding to characterize users' aesthetic preferences for images, and we believe that more rich explicit features are needed in modeling PIAA. Consequently, we propose an attribute-guided fine-grained feature-aware personalized image aesthetics assessment method, which can fully capture fine-grained features from multiple attributes to represent users' aesthetic preferences for images. To achieve this, we first build a fine-grained feature extraction (FFE) module to obtain the refined local features of image attributes to compensate for holistic features. The FFE module is then used to generate user-level features, which are combined with the image-level features to obtain user-preferred fine-grained feature representations. By training extensive users' PIAA tasks, the aesthetic distribution of most users can be transferred to the personalized scores of individual users. To enable our proposed model to learn more generalizable aesthetics among individual users, we incorporate the degree of dispersion between users' personalized scores and image aesthetic distribution as a coefficient in the loss function during model training. Experimental results on several PIAA databases show that our method outperforms existing mainstream PIAA methods, and can effectively infer users' personalized aesthetics of images.
Hancheng Zhu, Zhiwen Shao, Yong Zhou 0003, Guangcheng Wang, Pengfei Chen 0003, Leida Li
ACM Multimedia5
2023 Technical Quality-Assisted Image Aesthetics Quality Assessment
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Liwu Xu, Yuzhe Yang 0001
PRCV (11)3
2023 Dynamic Expert-Knowledge Ensemble for Generalizable Video Quality Assessment
abstract
Despite the impressive progress of supervised methods in quality assessment for in- the-wild videos, models trained on one domain often fail to generalize well to others due to the domain shifts caused by distortion diversity and content variation. Domain generalizable video quality assessment (VQA) methods that can work across domains remain an open research challenge. Although combining more data following the mixed-domain training strategy can improve the generalization performance to a certain extent, the specific knowledge from each source domain, which could potentially be useful for improving unseen domain generalization, is ignored in this principle. Motivated by this, we propose a domain generalizable VQA method named Dynamic Ensemble of Expert-Knowledge (DEEK), a novel framework that dynamically exploits the expert-knowledge from each source domain to achieve a generalizable ensemble prediction. Specifically, based on the multiple experts each trained to specialize in a particular source domain, we aim to exploit complementary information provided by the expert-knowledge. We effectively train an ensemble model by proposing a quality-sensitive InfoNCE loss to regularize the collaborative training of all experts in the contrastive learning formulation, aiming to exploit complementary information provided by the expert-knowledge when forming the ensemble. By dynamically integrating the experts according to their relevances to the target data, these expert-knowledge could be leveraged for better generalization. Experiments on five VQA datasets verify that our approach outperforms the state-of-the-arts by large margins.
Pengfei Chen 0003, Leida Li, Haoliang Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.1
2023 Image Aesthetics Assessment With Attribute-Assisted Multimodal Memory Network
abstract
Image aesthetics assessment (IAA) has attracted growing interest in recent years but is still challenging due to its highly abstract nature. Nowadays, more and more people tend to comment images shared on the social networks, which can provide rich aesthetics-aware semantic information from different aspects. Therefore, user comments of an image can be exploited as supplementary information for enhancing aesthetic representation learning. Previous researches have demonstrated that aesthetic attributes make significant effect on image aesthetic quality and humans’ aesthetic perception. Typically, people are used to give comments on an image from the perspective of aesthetic attributes, based on which the aesthetic quality of images can be inferred. Motivated by this, this paper presents an Attribute-assisted Multimodal Memory Network (AMM-Net) for image aesthetics assessment, which utilizes aesthetic attributes to model the interactions between visual and textual modalities. Specifically, we design two memory networks to capture the attribute-aware information most related to the image and associated comments respectively. Further, with multiple memory hops, attribute semantics shared by the two modalities are refined and cross-modal interactions are enhanced progressively. Finally, more discriminative aesthetic representations can be obtained for IAA. The experimental results and comparisons on two public multimodal IAA datasets demonstrate the superiority of the proposed model over the state-of-the-art methods. The source code is available athttps://github.com/zhutong0219/AMM-Net.
Leida Li, Tong Zhu 0003, Pengfei Chen 0003, Yuzhe Yang 0001, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.3
2022 Acknowledging the Unknown for Multi-label Learning with Single Positive Labels
Pengfei Chen 0003, Qiong Wang 0001, Guangyong Chen, Pheng-Ann Heng
ECCV (24)2
2022 Semantic Attribute Guided Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) measures the perceived beauty of images using a computational approach. People usually assess the aesthetics of an image according to semantic attributes, e.g., lighting, color, object emphasis, etc. However, the state-of-the-art IAA approaches usually follow the data-driven framework without considering the rich attributes contained in images. With this motivation, this paper presents a new semantic attribute guided IAA model, where the attention maps of semantic attributes are employed to enhance the representation ability of general aesthetic features for more effective aesthetics assessment. Specifically, we first design an attribute attention generation network to obtain the attention maps for different semantic attributes, which are utilized to weight the general aesthetic features, producing the semantic attribute-enhanced feature representations. Then, the Graph Convolutional Network (GCN) is employed to further investigate the inherent relationship among the enhanced aesthetic features, producing the final image aesthetics prediction. Extensive experiments and comparisons on three public IAA databases demonstrate the effectiveness of the proposed method.
Jiachen Duan, Pengfei Chen 0003, Leida Li, Jinjian Wu, Guangming Shi
VCIP2
2022 SPIQ: A Self-Supervised Pre-Trained Model for Image Quality Assessment
abstract
Blind image quality assessment (BIQA) has witnessed a flourishing progress due to the rapid advances in deep learning technique. The vast majority of prior BIQA methods try to leverage models pre-trained on ImageNet to mitigate the data shortage problem. These well-trained models, however, can be sub-optimal when applied to BIQA task that varies considerably from the image classification domain. To address this issue, we make the first attempt to leverage the plentiful unlabeled data to conduct self-supervised pre-training for BIQA task. Based on the distorted images generated from the high-quality samples using the designed distortion augmentation strategy, the proposed pre-training is implemented by a feature representation prediction task. Specifically, patch-wise feature representations corresponding to a certain grid are integrated to make prediction for the representation of the patch below it. The prediction quality is then evaluated using a contrastive loss to capture quality-aware information for BIQA task. Experimental results conducted on KADID-10 k and KonIQ-10 k databases demonstrate that the learned pre-trained model can significantly benefit the existing learning based IQA models.
Pengfei Chen 0003, Leida Li, Qingbo Wu 0001, Jinjian Wu
IEEE Signal Process. Lett.1
2022 Blind Image Quality Assessment for Authentic Distortions by Intermediary Enhancement and Iterative Training
abstract
With the boom of deep neural networks, blind image quality assessment (BIQA) has achieved great processes. However, the current BIQA metrics are limited when evaluating low-quality images as compared to medium-quality and high-quality images, which restricts their applications in real world problems. In this paper, we first identify that two challenges caused by distribution shift and long-tailed distribution lead to the compromised performance on low-quality images. Then, we propose an intermediary enhancement-based bilateral network with iterative training strategy for solving these two challenges. Drawing on the experience of transitive transfer learning, the proposed metric adaptively introduces enhanced intermediary images to transfer more information to low-quality images for mitigating the distribution shift. Our metric also adopts an iterative training strategy to deal with the long-tailed distribution. This strategy decouples feature extraction and score regression for better representation learning and regressor training. It not only transfers the knowledge learned from the earlier stage to the latter stage, but also makes the model pay more attention to long-tailed low-quality images. We conduct extensive experiments on five authentically distorted image quality datasets. The results show that our metric significantly improves the evaluating performance on low-quality images and delivers state-of-the-art intra-dataset results. During generalization tests, our metric also achieves the best cross-dataset performance.
Tianshu Song, Leida Li, Pengfei Chen 0003, Hantao Liu, Jiansheng Qian
IEEE Trans. Circuits Syst. Video Technol.3
2022 Contrastive Self-Supervised Pre-Training for Video Quality Assessment
abstract
Video quality assessment (VQA) task is an ongoing small sample learning problem due to the costly effort required for manual annotation. Since existing VQA datasets are of limited scale, prior research tries to leverage models pre-trained on ImageNet to mitigate this kind of shortage. Nonetheless, these well-trained models targeting on image classification task can be sub-optimal when applied on VQA data from a significantly different domain. In this paper, we make the first attempt to perform self-supervised pre-training for VQA task built upon contrastive learning method, targeting at exploiting the plentiful unlabeled video data to learn feature representation in a simple-yet-effective way. Specifically, we implement this idea by first generating distorted video samples with diverse distortion characteristics and visual contents based on the proposed distortion augmentation strategy. Afterwards, we conduct contrastive learning to capture quality-aware information by maximizing the agreement on feature representations of future frames and their corresponding predictions in the embedding space. In addition, we further introduce distortion prediction task as an additional learning objective to push the model towards discriminating different distortion categories of the input video. Solving these prediction tasks jointly with the contrastive learning not only provides stronger surrogate supervision signals, but also learns the shared knowledge among the prediction tasks. Extensive experiments demonstrate that our approach sets a new state-of-the-art in self-supervised learning for VQA task. Our results also underscore that the learned pre-trained model can significantly benefit the existing learning based VQA models. Source code is available at https://github.com/cpf0079/CSPT.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.1
2022 From Whole Video to Frames: Weakly-Supervised Domain Adaptive Continuous-Time QoE Evaluation
abstract
Due to the rapid increase in video traffic and relatively limited delivery infrastructure, end users often experience dynamically varying quality over time when viewing streaming videos. The user quality-of-experience (QoE) must be continuously monitored to deliver an optimized service. However, modern approaches for continuous-time video QoE estimation require densely annotating the continuous-time QoE labels, which is labor-intensive and time-consuming. To cope with such limitations, we propose a novel weakly-supervised domain adaptation approach for continuous-time QoE evaluation, by making use of a small amount of continuously labeled data in the source domain and abundant weakly-labeled data (only containing the retrospective QoE labels) in the target domain. Specifically, given a pair of videos from source and target domains, effective spatiotemporal segment-level feature representation is first learned by a combination of 2D and 3D convolutional networks. Then, a multi-task prediction framework is developed to simultaneously achieve continuous-time and retrospective QoE predictions, where a quality attentive adaptation approach is investigated to effectively alleviate the domain discrepancy without hampering the prediction performance. This approach is enabled by explicitly attending to the video-level discrimination and segment-level transferability in terms of the domain discrepancy. Experiments on benchmark databases demonstrate that the proposed method significantly improves the prediction performance under the cross-domain setting.
Leida Li, Pengfei Chen 0003, Weisi Lin, Mai Xu, Guangming Shi
IEEE Trans. Image Process.2
2022 Robust Medical Image Classification From Noisy Labeled Data With Global and Local Representation Guided Co-Training
abstract
Deep neural networks have achieved remarkable success in a wide variety of natural image and medical image computing tasks. However, these achievements indispensably rely on accurately annotated training data. If encountering some noisy-labeled images, the network training procedure would suffer from difficulties, leading to a sub-optimal classifier. This problem is even more severe in the medical image analysis field, as the annotation quality of medical images heavily relies on the expertise and experience of annotators. In this paper, we propose a novel collaborative training paradigm with global and local representation learning for robust medical image classification from noisy-labeled data to combat the lack of high quality annotated medical data. Specifically, we employ the self-ensemble model with a noisy label filter to efficiently select the clean and noisy samples. Then, the clean samples are trained by a collaborative training strategy to eliminate the disturbance from imperfect labeled samples. Notably, we further design a novel global and local representation learning scheme to implicitly regularize the networks to utilize noisy samples in a self-supervised manner. We evaluated our proposed robust learning strategy on four public medical image classification datasets with three types of label noise, i.e., random noise, computer-generated label noise, and inter-observer variability noise. Our method outperforms other learning from noisy label methods and we also conducted extensive experiments to analyze each component of our method.
Cheng Xue 0003, Lequan Yu, Pengfei Chen 0003, Qi Dou 0001, Pheng-Ann Heng
IEEE Trans. Medical Imaging3
2021 Beyond Class-Conditional Assumption: A Primary Attempt to Combat Instance-Dependent Label Noise
abstract
Supervised learning under label noise has seen numerous advances recently, while existing theoretical findings and empirical results broadly build up on the class-conditional noise (CCN) assumption that the noise is independent of input features given the true label. In this work, we present a theoretical hypothesis testing and prove that noise in real-world dataset is unlikely to be CCN, which confirms that label noise should depend on the instance and justifies the urgent need to go beyond the CCN assumption.The theoretical results motivate us to study the more general and practical-relevant instance-dependent noise (IDN). To stimulate the development of theory and methodology on IDN, we formalize an algorithm to generate controllable IDN and present both theoretical and empirical evidence to show that IDN is semantically meaningful and challenging. As a primary attempt to combat IDN, we present a tiny algorithm termed self-evolution average label (SEAL), which not only stands out under IDN with various noise fractions, but also improves the generalization on real-world noise benchmark Clothing1M. Our code is released. Notably, our theoretical analysis in Section 2 provides rigorous motivations for studying IDN, which is an important topic that deserves more research attention in future.
Pengfei Chen 0003, Junjie Ye 0002, Guangyong Chen, Pheng-Ann Heng
AAAI1
2021 Robustness of Accuracy Metric and its Inspirations in Learning with Noisy Labels
abstract
For multi-class classification under class-conditional label noise, we prove that the accuracy metric itself can be robust. We concretize this finding's inspiration in two essential aspects: training and validation, with which we address critical issues in learning with noisy labels. For training, we show that maximizing training accuracy on sufficiently many noisy samples yields an approximately optimal classifier. For validation, we prove that a noisy validation set is reliable, addressing the critical demand of model selection in scenarios like hyperparameter-tuning and early stopping. Previously, model selection using noisy validation samples has not been theoretically justified. We verify our theoretical results and additional claims with extensive experiments. We show characterizations of models trained with noisy labels, motivated by our theoretical results, and verify the utility of a noisy validation set by showing the impressive performance of a framework termed noisy best teacher and student (NTS). Our code is released.
Pengfei Chen 0003, Junjie Ye 0002, Guangyong Chen, Pheng-Ann Heng
AAAI1
2021 Foresee then Evaluate: Decomposing Value Estimation with Latent Future Prediction
abstract
Value function is the central notion of Reinforcement Learning (RL). Value estimation, especially with function approximation, can be challenging since it involves the stochasticity of environmental dynamics and reward signals that can be sparse and delayed in some cases. A typical model-free RL algorithm usually estimates the values of a policy by Temporal Difference (TD) or Monte Carlo (MC) algorithms directly from rewards, without explicitly taking dynamics into consideration. In this paper, we propose Value Decomposition with Future Prediction (VDFP), providing an explicit two-step understanding of the value estimation process: 1) first foresee the latent future, 2) and then evaluate it. We analytically decompose the value function into a latent future dynamics part and a policy-independent trajectory return part, inducing a way to model latent dynamics and returns separately in value estimation. Further, we derive a practical deep RL algorithm, consisting of a convolutional model to learn compact trajectory representation from past experiences, a conditional variational auto-encoder to predict the latent future dynamics and a convex return model that evaluates trajectory representation. In experiments, we empirically demonstrate the effectiveness of our approach for both off-policy and on-policy RL in several OpenAI Gym continuous control tasks as well as a few challenging variants with delayed reward.
Hongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen 0003, Chen Chen 0077, Yaodong Yang 0002, Luo Zhang 0002, Wulong Liu, Jianye Hao
AAAI4
2021 Unsupervised Curriculum Domain Adaptation for No-Reference Video Quality Assessment
abstract
During the last years, convolutional neural networks (C-NNs) have triumphed over video quality assessment (VQA) tasks. However, CNN-based approaches heavily rely on annotated data which are typically not available in VQA, leading to the difficulty of model generalization. Recent advances in domain adaptation technique makes it possible to adapt models trained on source data to unlabeled target data. However, due to the distortion diversity and content variation of the collected videos, the intrinsic subjectivity of VQA tasks hampers the adaptation performance. In this work, we propose a curriculum-style unsupervised domain adaptation to handle the cross-domain no-reference VQA problem. The proposed approach could be divided into two stages. In the first stage, we conduct an adaptation between source and target domains to predict the rating distribution for target samples, which can better reveal the subjective nature of VQA. From this adaptation, we split the data in target domain into confident and uncertain subdomains using the proposed uncertainty-based ranking function, through measuring their prediction confidences. In the second stage, by regarding samples in confident subdomain as the easy tasks in the curriculum, a fine-level adaptation is conducted between two subdomain-s to fine-tune the prediction model. Extensive experimental results on benchmark datasets highlight the superiority of the proposed method over the competing methods in both accuracy and speed. The source code is released at https://github.com/cpf0079/UCDA.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
ICCV1
2021 Noise against noise: stochastic label noise helps combat inherent label noise
Pengfei Chen 0003, Guangyong Chen, Junjie Ye 0002, Pheng-Ann Heng
ICLR1
2021 Temporal Reasoning Guided QoE Evaluation for Mobile Live Video Broadcasting
abstract
Quality of experience (QoE) that serves as a direct evaluation of viewing experience from the end users is of vital importance for network optimization, and should be constantly monitored. Unlike existing video-on-demand streaming services, real-time interactivity is critical to the mobile live broadcasting experience for both broadcasters and their audiences. While existing QoE metrics that are validated on limited video contents and synthetic stall patterns have shown effectiveness in their trained QoE benchmarks, a common caveat is that they often encounter challenges in practical live broadcasting scenarios, where one needs to accurately understand the activity in the video with fluctuating QoE and figure out what is going to happen to support the real-time feedback to the broadcaster. In this paper, we propose a temporal relational reasoning guided QoE evaluation approach for mobile live video broadcasting, namely TRR-QoE, which explicitly attends to the temporal relationships between consecutive frames to achieve a more comprehensive understanding of the distortion-aware variation. In our design, video frames are first processed by deep neural network (DNN) to extract quality-indicative features. Afterwards, besides explicitly integrating features of individual frames to account for the spatial distortion information, multi-scale temporal relational information corresponding to diverse temporal resolutions are made full use of to capture temporal-distortion-aware variation. As a result, the overall QoE prediction could be derived by combining both aspects. The results of experiments conducted on a number of benchmark databases demonstrate the superiority of TRR-QoE over the representative state-of-the-art metrics.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Yabin Zhang 0002, Weisi Lin
IEEE Trans. Image Process.1
2020 RIRNet: Recurrent-In-Recurrent Network for Video Quality Assessment
abstract
Video quality assessment (VQA), which is capable of automatically predicting the perceptual quality of source videos especially when reference information is not available, has become a major concern for video service providers due to the growing demand for video quality of experience (QoE) by end users. While significant advances have been achieved from the recent deep learning techniques, they often lead to misleading results in VQA tasks given their limitations on describing 3D spatio-temporal regularities using only fixed temporal frequency. Partially inspired by psychophysical and vision science studies revealing the speed tuning property of neurons in visual cortex when performing motion perception (i.e., sensitive to different temporal frequencies), we propose a novel no-reference (NR) VQA framework named Recurrent-In-Recurrent Network (RIRNet) to incorporate this characteristic to prompt an accurate representation of motion perception in VQA task. By fusing motion information derived from different temporal frequencies in a more efficient way, the resulting temporal modeling scheme is formulated to quantify the temporal motion effect via a hierarchical distortion description. It is found that the proposed framework is in closer agreement with quality perception of the distorted videos since it integrates concepts from motion perception in human visual system (HVS), which is manifested in the designed network structure composed of low- and high- level processing. A holistic validation of our methods on four challenging video quality databases demonstrates the superior performances over the state-of-the-art methods.
Pengfei Chen 0003, Leida Li, Lei Ma 0003, Jinjian Wu, Guangming Shi
ACM Multimedia1
2019 QoE Evaluation for Live Broadcasting Video
abstract
The great variations of videographic skills in shot environment, photographic apparatus, compression and processing protocols give rise to very complicated impairments in the live broadcasting videos, which can adversely impact the quality of experience (QoE) of end users. Evaluating QoE of these videos is of great significance. Given the fact that there is still no publicly available database that studies the combined effects of the distortions in the live broadcasting videos, we have built the Live Broadcasting Video Database (LBVD) with the associated QoE scores. Towards depicting the distortions in live broadcasting videos, totally 1013 videos were included in the database, with diversified, authentic distortions. A subjective evaluation of these videos is conducted, and the correlation results between the tested state-of-the-art objective metrics and the subjective QoE scores on this database reveal that further studies are in urgent need for a better objective QoE metric dedicated to the live broadcasting videos.
Pengfei Chen 0003, Leida Li, Yipo Huang, Fengfeng Tan
ICIP1
2019 Understanding and Utilizing Deep Neural Networks Trained with Noisy Labels
abstract
Noisy labels are ubiquitous in real-world datasets, which poses a challenge for robustly training deep neural networks (DNNs) as DNNs usually have the high capacity to memorize the noisy labels. In this paper, we find that the test accuracy can be quantitatively characterized in terms of the noise ratio in datasets. In particular, the test accuracy is a quadratic function of the noise ratio in the case of symmetric noise, which explains the experimental findings previously published. Based on our analysis, we apply cross-validation to randomly split noisy datasets, which identifies most samples that have correct labels. Then we adopt the Co-teaching strategy which takes full advantage of the identified samples to train DNNs robustly against noisy labels. Compared with extensive state-of-the-art methods, our strategy consistently improves the generalization performance of DNNs under both synthetic and real-world training noise.
Pengfei Chen 0003, Benben Liao, Guangyong Chen, Shengyu Zhang 0002
ICML1
2019 Blind quality index for tone-mapped images based on luminance partition
Pengfei Chen 0003, Leida Li, Xinfeng Zhang 0001, Shanshe Wang, Allen Tan
Pattern Recognit.1