Yipo Huang

dblp:254/8276 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-0908-2180ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning
abstract
Livestreaming has become increasingly prevalent in modern visual communication, where automatic camera quality tuning is essential for delivering superior user Quality of Experience (QoE). Such tuning requires accurate blind image quality assessment (BIQA) to guide parameter optimization decisions. Unfortunately, the existing BIQA models typically only predict an overall coarse-grained quality score, which cannot provide fine-grained perceptual guidance for precise camera parameter tuning. To bridge this gap, we first establish FGLive-10K, a comprehensive fine-grained BIQA database containing 10,185 high-resolution images captured under varying camera parameter configurations across diverse livestreaming scenarios. The dataset features 50,925 multi-attribute quality annotations and 19,234 fine-grained pairwise preference annotations. Based on FGLive-10K, we further develop TuningIQA, a fine-grained BIQA metric for livestreaming camera tuning, which integrates human-aware feature extraction and graph-based camera parameter fusion. Extensive experiments and comparisons demonstrate that TuningIQA significantly outperforms state-of-the-art BIQA methods in both score regression and fine-grained quality ranking, achieving superior performance when deployed for livestreaming camera tuning.
Xiangfei Sheng, Zhichao Duan 0002, Xiaofeng Pan, Yipo Huang, Zhichao Yang 0013, Pengfei Chen 0003, Leida Li
AAAI4
2026 HumanCrop-Thinker: An inference-driven framework with explicit thinking for explainable human-centric image cropping
Yipo Huang, Pengfei Chen 0003, Leida Li
Expert Syst. Appl.2
2026 Learning Scene-Invariant Distribution for Generalizable Blind Image Quality Assessment
abstract
The inherent diversity of visual scenes poses a fundamental challenge in blind image quality assessment (BIQA), which has become a major obstacle to the model generalization. In this study, we found that human annotations for images with different visual scenes exhibit distinct quality distribution discrepancies. The existing BIQA models tend to overfit to such diversified distributions, which in turn leads to compromised model generalizability, especially when dealing with unseen scenes in the real-world scenario. Motivated by the above facts, this paper presents a generalizable BIQA model by learning Scene-INvariant Distribution, named SIND. Specifically, we propose a distribution alignment framework to alleviate the distribution discrepancy for quality regression models, which is achieved by automatically scaling and shifting the cross-scene distributions into a unified distribution. Then, the aligned unified distribution is leveraged to supervise the model training, achieving scene-invariant and quality-aware feature representation. In addition, a token-complementary patch reasoning network is designed to extract comprehensive quality-aware features from both the image overview and detail, achieving more accurate quality prediction. Extensive experiments for both image technical- and aesthetic-quality assessment tasks show the superiority of the proposed SIND model over the state-of-the-arts. Moreover, the proposed framework is model-agnostic and can enhance model generalizability without incurring extra inference costs. The proposed method won the championship in the NTIRE 2024 Portrait Quality Assessment Challenge. Codes will be available at https://github.com/ZachL1/SIND.
Yipo Huang, Zhichao Duan 0002, Pengfei Chen 0003, Leida Li, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1
2025 VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment: Methods and Results
abstract
This paper reviews the VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment. The competition aims to foster innovation in both subjective and objective VQA techniques tailored to live broadcasting videos, addressing the unique challenges posed by live streaming impairments while emphasizing the evaluation of QoE. The grand challenge used live broadcasting database LBVD which consists of 1013 videos focusing on distortion in live broadcasting videos. The competition had 14 participants and 5 teams submitted valid solutions for the final testing phase. The proposed solutions have shown significant progress in areas such as combining traditional feature engineering with deep learning models, achieved state-of-the-art performances for LBVD. Team ATHENA-Live-QoE and Team HZX Force tied for the first position. The dataset can be found at https://github.com/cpf0079/LBVD.
Wenqi Fei, Yuhua Zhang, MohammadAli Hamidi, Hadi Amirpour, Erjia Xiao, Zhenjie Su, Hao Cheng 0015, Yu Liu 0023, Wei Zhou 0021, Yanbiao Ma, Renjing Xu, Long Chen 0015, Xiaoshuai Hao, Yipo Huang, Tushar Shinde
VCIP17
2025 Multi-Modality Multi-Attribute Contrastive Pre-Training for Image Aesthetics Computing
abstract
In the Image Aesthetics Computing (IAC) field, most prior methods leveraged the off-the-shelf backbones pre-trained on the large-scale ImageNet database. While these pre-trained backbones have achieved notable success, they often overemphasize object-level semantics and fail to capture the high-level concepts of image aesthetics, which may only achieve suboptimal performances. To tackle this long-neglected problem, we propose a multi-modality multi-attribute contrastive pre-training framework, targeting at constructing an alternative to ImageNet-based pre-training for IAC. Specifically, the proposed framework consists of two main aspects. 1) We build a multi-attribute image description database with human feedback, leveraging the competent image understanding capability of the multi-modality large language model to generate rich aesthetic descriptions. 2) To better adapt models to aesthetic computing tasks, we integrate the image-based visual features with the attribute-based text features, and map the integrated features into different embedding spaces, based on which the multi-attribute contrastive learning is proposed for obtaining more comprehensive aesthetic representation. To alleviate the distribution shift encountered when transitioning from the general visual domain to the aesthetic domain, we further propose a semantic affinity loss to restrain the content information and enhance model generalization. Extensive experiments demonstrate that the proposed framework sets new state-of-the-arts for IAC tasks.
Yipo Huang, Leida Li, Pengfei Chen 0003, Haoning Wu 0001, Weisi Lin, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception
abstract
The highly abstract nature of image aesthetics perception (IAP) poses a significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma, resulting in MLLMs falling short of aesthetics perception capabilities. To address the above challenge, we first introduce a comprehensively annotated Aesthetic Multi-Modality Instruction Tuning (AesMMIT) dataset, which serves as the footstone for building multi-modality aesthetics foundation models. Specifically, to align MLLMs with human aesthetics perception, we construct a corpus-rich aesthetic critique database with 21,904 diverse-sourced images and 88K human natural language feedbacks, which are collected via progressive questions, ranging from coarse-grained aesthetic grades to fine-grained aesthetic descriptions. To ensure that MLLMs can handle diverse queries, we further prompt GPT to refine the aesthetic critiques and assemble the large-scale aesthetic instruction tuning dataset, i.e. AesMMIT, which consists of 409K multi-typed instructions to activate stronger aesthetic capabilities. Based on the AesMMIT database, we fine-tune the open-sourced general foundation models, achieving multi-modality Aesthetic Expert models, dubbed AesExpert. Extensive experiments demonstrate that the proposed AesExpert models deliver significantly better aesthetic perception performances than the state-of-the-art MLLMs, including the most advanced GPT-4V and Gemini-Pro-Vision. Project Page: https://yipoh.github.io/aes-expert/.
Yipo Huang, Xiangfei Sheng, Zhichao Yang 0013, Zhichao Duan 0002, Pengfei Chen 0003, Leida Li, Weisi Lin, Guangming Shi
ACM Multimedia1
2024 Coarse-to-Fine Image Aesthetics Assessment With Dynamic Attribute Selection
abstract
Image aesthetics assessment (IAA) is an interesting but challenging task, owing to the ineffable nature of human sense of beauty. The study of IAA has evolved from simple binary classification to more complex score regression and distribution prediction. It is effortless for people to perform aesthetic binary classification,i.e., aesthetically pleasing or not. However, further judgment on the fine-level scalar aesthetic score is complex and typically determined by aesthetic attributes presented in the image, such as content, lighting and color. Motivated by the above facts, this paper presents a Coarse-to-fine image Aesthetics assessment model guided by Dynamic Attribute Selection, dubbed CADAS. The underlying idea is to simulate the process of human aesthetic perception by performing coarse-to-fine aesthetic reasoning. Specifically, a hierarchical AttributeNet is first pre-trained by imitating the staged mechanism of human aesthetic experience, producing the candidate aesthetic attributes. Then, an AestheticNet is introduced to perform the coarse-level binary classification, based on which a confidence-based attribute selection strategy is designed to dynamically pick out the dominant aesthetic attributes from the candidate ones. Finally, a self-attention-based FusionNet is designed to explore the interaction between dominant aesthetic attributes and aesthetic features, producing the fine-level aesthetic prediction. Extensive experiments demonstrate that the proposed model is superior to the state-of-the-arts. Furthermore, CADAS is also able to output the dominant aesthetic attributes in images, facilitating model explainability.
Yipo Huang, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Guangming Shi
IEEE Trans. Multim.1
2023 Quality Prediction of View Synthesis Based on Curriculum-Style Structure Generation
abstract
Existing quality metrics of view synthesis usually perform on synthesized images, which are produced based on a computationally expensive depth-image-based rendering (DIBR) process. Moreover, current metrics quantify quality by extracting hand-crafted features, which may fail to fully capture the complex distortion characteristics. With the success of deep learning on numerous computer vision tasks, it has become possible to utilize convolutional neural networks to predict the quality of DIBR-synthesized images. In this letter, we propose a deep model to predict the quality of view synthesis based on Curriculum-style Structure Generation without conducting the DIBR process. Specifically, considering that the distortion of view synthesis is mainly manifested in the destruction of image structure, a structure generation network is first built to learn the structure of the new view from the original one by curriculum-style training. Then, we transfer the prior knowledge learned from the last phase into the quality prediction network for measuring the structure distortion, based on which a regressor is introduced to produce the quality score. Experimental results prove the advantages of the proposed model.
Haozhi Shi, Yipo Huang, Lanmei Wang
IEEE Signal Process. Lett.2
2023 Theme-Aware Visual Attribute Reasoning for Image Aesthetics Assessment
abstract
People usually assess image aesthetics according to visual attributes, e.g., interesting content, good lighting and vivid color, etc. Further, the perception of visual attributes depends on the image theme. Therefore, the inherent relationship between visual attributes and image theme is crucial for image aesthetics assessment (IAA), which has not been comprehensively investigated. With this motivation, this paper presents a new IAA model based on Theme-Aware Visual Attribute Reasoning (TAVAR). The underlying idea is to simulate the process of human perception in image aesthetics by performing bilevel reasoning. Specifically, a visual attribute analysis network and a theme understanding network are first pre-trained to extract aesthetic attribute features and theme features, respectively. Then, the first level Attribute-Theme Graph (ATG) is built to investigate the coupling relationship between visual attributes and image theme. Further, a flexible aesthetics network is introduced to extract general aesthetic features, based on which we built the second level Attribute-Aesthetics Graph (AAG) to mine the relationship between theme-aware visual attributes and aesthetic features, producing the final aesthetic prediction. Extensive experiments on four public IAA databases demonstrate the superiority of the proposed TAVAR model over the state-of-the-arts. Furthermore, TAVAR features better explainability due to the use of visual attributes.
Leida Li, Yipo Huang, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.2
2023 Explainable and Generalizable Blind Image Quality Assessment via Semantic Attribute Reasoning
abstract
Blind image quality assessment (BIQA) that can directly evaluate image quality without perfect-quality reference has been a long-standing research topic. Although the existing BIQA models have achieved very encouraging performance, the lack of explainability and generalization ability limits their real-world applications to a great extent. People usually assess image quality according to semantic attributes, e.g., brightness, color, contrast, noise and sharpness. Furthermore, judgment on image quality is also impacted by the scene presented in the image. Therefore, the inherent relationship between semantic attributes and scenes is crucial for image quality assessment, which has rarely been explored yet. With this motivation, this paper presents a Semantic Attribute Reasoning based image QUality Evaluator (SARQUE). Specifically, we propose a two-stream network to predict semantic attributes and scene categories from distorted images. To investigate the inherent relationship between the semantic attributes and scene category, a semantic reasoning module is further proposed based on the graph convolution network (GCN), producing the final quality score. Extensive experiments conducted on five in-the-wild image quality databases demonstrate the superiority of the proposed SARQUE model over the state-of-the-arts. Furthermore, the proposed model features better explainability and generalization ability due to the use of semantic attributes.
Yipo Huang, Leida Li, Yuzhe Yang 0001, Yandong Guo
IEEE Trans. Multim.1
2021 Predicting the Quality of View Synthesis With Color-Depth Image Fusion
abstract
With the increasing prevalence of free-viewpoint video applications, virtual view synthesis has attracted extensive attention. In view synthesis, a new viewpoint is generated from the input color and depth images with a depth-image-based rendering (DIBR) algorithm. Current quality evaluation models for view synthesis typically operate on the synthesized images, i.e. after the DIBR process, which is computationally expensive. So a natural question is that can we infer the quality of DIBR-based synthesized images using the input color and depth images directly without performing the intricate DIBR operation. With this motivation, this paper presents a no-reference image quality prediction model for view synthesis via COlor-Depth Image Fusion, dubbed CODIF, where the actual DIBR is not needed. First, object boundary regions are detected from the color image, and a Wavelet-based image fusion method is proposed to imitate the interaction between color and depth images during the DIBR process. Then statistical features of the interactional regions and natural regions are extracted from the fused color-depth image to portray the influences of distortions in color/depth images on the quality of synthesized views. Finally, all statistical features are utilized to learn the quality prediction model for view synthesis. Extensive experiments on public view synthesis databases demonstrate the advantages of the proposed metric in predicting the quality of view synthesis, and it even suppresses the state-of-the-art post-DIBR view synthesis quality metrics.
Leida Li, Yipo Huang, Jinjian Wu, Ke Gu 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 No-reference quality assessment for live broadcasting videos in temporal and spatial domains
abstract
Nowadays, live broadcasting video has become increasingly popular and high‐quality live broadcasting video is highly needed. In practice, live broadcasting videos usually undergo several processing stages, which inevitably introduce multiple distortions, e. g. frame freezing and intensity mutation, causing the degraded quality of experience. However, little work has been done to the quality evaluation of live broadcasting videos, which may hinder the further development of more advanced live broadcasting video delivery systems. Motivated by this, this study presents a no‐reference quality evaluation model for live broadcasting videos (LBVQA) in temporal and spatial domains. In the temporal domain, statistic features are extracted to measure the frame freezing and intensity mutation, and the entropy‐based feature is extracted to describe the global jitter. In the spatial domain, blurring is measured based on phase coherence, and abnormal exposure ratio is calculated based on an adaptive threshold. Finally, all features are fed into a backpropagation neural network to train the quality prediction model. Experimental results on the Live Broadcasting Video Database demonstrate the advantages of the proposed metric over the state‐of‐the‐art image and video quality metrics.
Yipo Huang, Leida Li, Yu Zhou 0009, Bo Hu 0008
IET Image Process.1
2020 Blind Quality Index of Depth Images Based on Structural Statistics for View Synthesis
abstract
The quality of depth images is crucial for virtual view synthesis. However, the quality assessment of depth images is still largely unexplored. This letter presents a blind quality metric of Depth image based on Structural Statistics (DSS). The design philosophy is inspired by the fact that structural distortion in the depth images usually leads to geometric distortion, which is the main cause for degraded quality of synthesized views. Specifically, the statistical features for shape and orientation are calculated based on discrete orthogonal moments and gradients, generating two groups of quality-aware features. Then, the quality model is built from the extracted statistical features using a regression module. The experimental results demonstrate the effectiveness of the proposed metric.
Yipo Huang, Leida Li, Hancheng Zhu, Bo Hu 0008
IEEE Signal Process. Lett.1
2019 QoE Evaluation for Live Broadcasting Video
abstract
The great variations of videographic skills in shot environment, photographic apparatus, compression and processing protocols give rise to very complicated impairments in the live broadcasting videos, which can adversely impact the quality of experience (QoE) of end users. Evaluating QoE of these videos is of great significance. Given the fact that there is still no publicly available database that studies the combined effects of the distortions in the live broadcasting videos, we have built the Live Broadcasting Video Database (LBVD) with the associated QoE scores. Towards depicting the distortions in live broadcasting videos, totally 1013 videos were included in the database, with diversified, authentic distortions. A subjective evaluation of these videos is conducted, and the correlation results between the tested state-of-the-art objective metrics and the subjective QoE scores on this database reveal that further studies are in urgent need for a better objective QoE metric dedicated to the live broadcasting videos.
Pengfei Chen 0003, Leida Li, Yipo Huang, Fengfeng Tan
ICIP3